Skip to content

Keep socket discovery and remote forwarding on this user's socket - #15151

Closed
austinywang wants to merge 168 commits into
local-socket-peer-checksfrom
socket-discovery-owner-checks
Closed

austinywang wants to merge 168 commits into
local-socket-peer-checksfrom
socket-discovery-owner-checks

Conversation

@austinywang

@austinywang austinywang commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #15142. It's based on local-socket-peer-checks, so review and merge that first. This PR closes two socket paths #15142 left open.

Summary

Socket discovery. To find the app, the CLI probes candidate control sockets, and some of them sit at fixed names in the shared /tmp directory. Before this change the probe checked who owned the socket file, then accepted any listener that took the connection. It now also reads the listener's user with #15142's UnixSocketPeerCheck after the connect finishes. A socket another user is listening on fails the probe, so discovery moves on to the next candidate. The probe moved out of CLISocketPathResolver into CmuxFoundation as UnixSocketConnectProbe so it can be tested. The existing lstat owner check on the socket file stays where it was. All CLISocketPathResolver discovery goes through this one probe.

Remote forwarding target. workspace.remote.configure takes a local_socket_path parameter. The app stored whatever the client sent, and a remote workspace's CLI relay and cloud CLI bridge then forwarded remote commands to that path. The app now stores its own control socket path (currentSocketPathForRemoteRestore()) and ignores the value the client sends. A non-blank local_socket_path still turns forwarding on, and a missing or blank one leaves it off. The CLI already sends the socket it connected through, which is this app's socket, so normal cmux ssh and cloud flows behave the same. One small side effect: a blank value used to be stored as an empty string, which the fork-conversation menu treated as a reachable socket. It's now stored as nil.

Relay policy. No change. workspace.remote.configure isn't in RemoteRelayRoutingSchema, so RemoteRelayCommandPolicy already refuses it when it comes through the remote relay. Only local socket clients can call it, and they now can't choose where forwarded commands go.

Testing

Regression tests were committed first, in beb4c14, and failed there. The fix is 6ce2e23.

Command beb4c14 (tests only) 6ce2e23 (fix)
swift test --disable-index-store --package-path Packages/macOS/CmuxFoundation --filter UnixSocketConnectProbeTests 3 tests, 1 failed: refusesAListenerRunningAsAnotherUser 3 tests passed
swift test --disable-index-store --package-path Packages/macOS/CmuxControlSocket --filter ControlWorkspaceRemoteLocalSocketPathTests 3 tests (9 cases) failed, 8 issues 3 tests (9 cases) passed

The probe test simulates another user by expecting a user ID other than the current one; there was no second local account to run a real listener. python3 scripts/verify-local.py --affected origin/main --swift-changed origin/main passed Swift syntax, package groups and feature flags.

Not verified: the cmux-cli and app targets and cmuxTests weren't compiled or run locally. That covers the one-line CLISocketPathResolver change and the TerminalController+ControlWorkspaceContext call site. There was no dogfood in a tagged build, so cmux ssh and cloud CLI forwarding weren't checked live.

Follow-up not included. Moving tagged and debug sockets from /tmp into a per-user directory. That touches socket path resolution in the app, CLI, scripts, docs and tests, which makes it too big to add here safely. With #15142 and this PR, the app and CLI refuse a socket another user holds at one of those names, so what's left is that user blocking the name, not intercepting commands.

Changelog

Fixed: The CLI skips control sockets another user is listening on, and remote workspaces always forward commands to the app's own socket

Checklist

  • Behavior changes have added or updated tests, or Testing says why not
  • No UI, settings, menu, schema, help-text or user-facing docs change
  • No v2 socket method allowlisted for cmux ssh
  • Reviewed with a subagent before merge (cmux-review), and all bot and human review comments resolved

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.


Summary by cubic

Closes two socket paths left open: the CLI now checks the listener's user when probing candidate control sockets, and remote workspaces forward commands to the app's own socket instead of the path the client sends.

Socket discovery

  • The probe moved from CLISocketPathResolver into UnixSocketConnectProbe in CmuxFoundation and now checks the listener's user with UnixSocketPeerCheck after connecting.
  • A socket another user is listening on fails the probe, so discovery moves to the next candidate; the existing lstat file-owner check stays.

Remote forwarding target

  • workspace.remote.configure stores currentSocketPathForRemoteRestore() and ignores the client's local_socket_path value, which only turns forwarding on or off.
  • A blank value now stores nil instead of an empty string, so the fork-conversation menu no longer treats it as a reachable socket.

Regression tests were committed first and failed there; the fix commits followed. The remaining commits are merges from main and don't change these two behaviors.

Written for commit 7792b9a. Summary will update on new commits.

Review in cubic

lawrencecchen and others added 10 commits September 27, 2026 20:36
…lures (#15103)

* test: CLI Sentry filter must treat rate_limited, auth_required and socket EPERM as caller state

Covers Sentry CMUXTERM-MACOS-3JFC (rate_limited), CMUXTERM-MACOS-3JNR
(auth_required) and CMUXTERM-MACOS-3JHJ (socket connect EPERM per hook).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: stop CLI Sentry floods from caller state and unattributed journal failures

- rate_limited, auth_required and auth_failed protocol replies are caller
  state, not server failures (CMUXTERM-MACOS-3JFC, CMUXTERM-MACOS-3JNR).
- Socket-connect EPERM without trusted sandbox provenance stays reportable
  but groups as socket-connect-denied and is throttled like command
  timeouts (CMUXTERM-MACOS-3JHJ).
- Agent journal append failures now pass the underlying error and a stable
  failure kind, so lifecycle/socket-missing states are classified and real
  rejections say why (CMUXTERM-MACOS-3JFH reported every failure as
  unresolved-target).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Bound both simulator launch setup and report notification waiters to the configured deadline.
…15124)

* ci: charge newer runs one root runner each when gui runners are on

The picker charged a newer run's whole marker peak on the std pool to its
root runners. With gui runners, shards, lag and cli-product take the gui
label and side lanes never hold a root runner, so a newer run holds only
its admission there. Three newer runs with a 32-machine peak left 0 of 15
root runners for a run while they held 3, and its admission and follow-on
jobs went to Blacksmith (run 36371179217).

choose() now passes the newer runs' run count as their root charge when
the slots name gui runners; without gui runners the whole peak stays the
charge, since a marker does not split side lanes from root jobs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: read gui runners per pool for the root charge

Review: a pool's newer runs hold one root runner only where that pool's
gui label has a slot count (gui_runner()); any other pool keeps the peak.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…es (#15122)

* Document and tool in-place cmux-tui upgrades for running Cloud machines

Machines never update themselves, so a contract change strands every
running machine. #14125 refused attach for 16 running pre-snapshot-v2
machines of 10 users; most already ran a compatible daemon.

- docs/cloud-guest-upgrades.md: what can reach a running machine, the
  compatibility rules for cmux-tui and web changes, the upgrade runbook
  with the guarded backfill, and the plan for machines that cannot be
  upgraded.
- web/services/vms/images/devbox/cmux-tui-upgrade: the guest script used
  on 2026-09-28 to upgrade 20 machines with no terminal lost. It skips
  machines whose supervisor predates trusted carrier, waits for slow
  daemon starts, and rolls back a daemon that crashes or never listens.
- web/scripts/upgrade-fleet-cmux-tui.ts: runs it with the bake's pinned
  install command, targeting the default image's cmux-tui build.
- web/AGENTS.md, cmux-tui/AGENTS.md: point future changes at the rules.
- freestyle.ts and the devbox README: drop the claims that no older rows
  exist and that attach heals pin drift.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Move the guest upgrade script out of the baked image directory

vm-devbox-image.test.ts pins the devbox directory to the files the bake
ships; the upgrade script is fleet tooling, not image content.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Harden the guest upgrade script per review

- One run directory per run and a per-machine flock: concurrent runs
  never share an install command or result (second run: SKIP busy).
- Install nothing unless the running daemon listens.
- Save the replaced binary per run, verified by sha256.
- A lost terminal host or a lower terminal count is FAIL, not OK.
- Rollback verifies the restored binary and trusts it only once the old
  daemon serves again; otherwise FAIL rollback-daemon-unhealthy, since
  the new daemon may have migrated on-disk state.
- Runner rejects malformed machine ids instead of dropping them.
- Backfill accepts a machine that recorded only IPv6.

Verified live on vm-868f46: two concurrent runs gave SKIP busy and
OK upgraded terminals=4->4 with every host alive.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Report an unknown terminal count as UNVERIFIED, not OK

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…5101)

* test: expect surface_unavailable to be a routine CLI outcome

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: report a non-running terminal as surface_unavailable in read_text

surface.read_text returned internal_error "Failed to read terminal text"
whenever the resolved terminal had no live Ghostty surface: a hibernated
agent, a restore awaiting admission, or a start that missed the deadline.
Those are surface states, so reply surface_unavailable with a localized,
actionable message and the surface id, and treat surface_unavailable as a
routine CLI protocol outcome for Sentry (CMUXTERM-MACOS-3JFD).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Translate socket.terminal.notRunning for all catalog locales

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Raise docs search contrast

The search field used code-bg at 60% opacity with a transparent border, so
in light mode it nearly vanished into the page, and the placeholder and icon
used muted at 40% (about 1.7:1). Use the full code-bg fill with a border-border
outline and full-strength muted for the placeholder, icon, and result metadata.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Darken light-theme muted text to meet 4.5:1

#737373 on the #f5f5f5 search field is 4.35:1, below WCAG AA for the
11-13px text there. #707070 reaches 4.54:1; dark theme is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Use full-strength muted for the docs search focus border

The 60% border was about 2.3:1 on the light field, below the 3:1 non-text
target, and close to the hover border.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(cloud): cover foreground recovery state

The new behavior cases intentionally fail on current main: automatic refresh keeps the stale machine-list failure visible and app activation starts no recovery read.\n\nRefs #15100

* fix(cloud): recover machine list on foreground activation

Mark automatic list reads as recovery at the request owner so stale transient failures are neutral while the read is in flight. Keep routine polls actionable, and observe app foreground activation alongside existing wake and network recovery.\n\nFixes #15100

* fix(cloud): keep refresh setter access scoped

* fix(cloud): keep refresh mutations at the model owner
* test: require silent Cloud sidebar drags

* fix: suppress Cloud sidebar drag hints and redraw work
…socket

Move the CLI's socket discovery probe into CmuxFoundation as
UnixSocketConnectProbe and route workspace.remote.configure's
local_socket_path through ControlWorkspaceRemoteLocalSocketPath, both
unchanged in behavior. The new tests expect the probe to refuse a listener
running as another user and the forwarding path to be the app's own control
socket; they fail here.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Socket discovery now checks the connected listener's user with
UnixSocketPeerCheck, on top of the existing socket file owner check, so the
CLI skips a listener another user runs. workspace.remote.configure stores the
app's own control socket for the remote CLI relay and cloud CLI bridge
instead of the path the client sends; the client's value only turns
forwarding on.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository: manaflow-ai/cmux/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: 3fa2a1c8-9e84-412b-bc43-d23dd41ad881

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@blacksmith-sh

This comment has been minimized.

@github-actions

Copy link
Copy Markdown
Contributor

CI failure attribution

CI failed on 6ce2e232a6 (run 36377920548 attempt 1): 1 unknown.

Job Verdict Why
guards / workflow-guard-tests / release-ios unknown no known signature; failed step: Validate iOS package conventions for this change

Not re-run automatically: guards / workflow-guard-tests / release-ios is not a machine failure.

Written by scripts/ci/classify_failures.py (ci-failure-attribution.yml); signatures are its SIGNATURES table. A machine verdict is the runner's fault, not this PR's.

austinywang and others added 17 commits September 27, 2026 21:49
…ed agents (#14420)

* test: distinguish Mac discoverability from discovery and identity failures

* fix: require Mac discovery consent and report confirmed opt-out

* test: cover Mac terminal grid updates after coalesced render ticks

* fix: propagate Mac terminal grid changes without reopening mirrors

* fix: make oversized Mac mirrors locally scrollable

* fix: adopt device split reservations in their target pane

* fix: include prose wake driver in macOS target

* test: bound distinct device grid queue events

* fix: harden device layout and grid delivery

* fix: avoid retaining device mirror attachment

* fix: compile device pane reconciliation

* fix: use native pane identifiers in device projection

* fix: expose reservation identity binding

* fix: use reservation parameter in device materialization

* fix: bound grid queue ownership and pending layout admission

* fix: avoid namespace policy lint violation

* Hand a reserved pane's input to an adopting device mirror safely

A device mirror that adopts an optimistic reserved pane now takes over the
pane's input relay with its own byte router, which fixes the macOS compile
error where the provider passed a DeviceTerminalInputRouter to a relay that
only accepted the Cloud router.

- Reserved panes for devices install no named-key resolver. The device router
  sends bytes only, so a resolver would drop Enter, arrows and Tab.
- The relay is handed over only when adoption succeeded, and only on an attach
  that is not immediately replaced by a queued replay, so input typed before
  the terminal attached is delivered in order and not dropped mid-handoff.
- Input typed while the source Mac is unreachable is discarded on detach and
  never replays after reconnecting, matching panes the router created itself.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Import Bonsplit where the layout projection test names a pane

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that a replacement grid overflow reaches the drain

A Mac grid replacement larger than the byte budget returns overflow
without recording it, so the fan-out path never closes the connection.
A running drain can also finish over a pending overflow and strand it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Close the connection on every event queue overflow

The queue's pending overflow flag is the one signal a drain uses to close
the connection. Every overflow result now records it and claims the drain
through one helper, a running drain checks it on each pass, and neither
finishDrain nor claimDrain lets an unconsumed overflow go unobserved. The
unrecorded .overflow constant is removed so no branch can skip the flag.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that two device splits adopt their own reserved panes

Two outstanding device splits in one remote workspace must each adopt
the reservation bound to the terminal it created, and a terminal that
no request created must never take an unbound reservation's pane or
queued input. The existing split test now binds its reservation through
the create receipt and uses a UUID remote workspace, which the layout
coordinator requires.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Adopt a device split's pane only for the terminal bound to it

Device layout reconciliation took the first pending reservation in the
remote workspace and adopted it for any new terminal whose id matched
the reservation's source placement. Before its create receipt binds,
that placement names the split source, so a second outstanding split,
or a terminal no request created, could land in the wrong pane and
receive another terminal's queued input.

Invariant: a reservation's pane is used only for the terminal its
create receipt bound (`boundResourceID`). `cloudPendingCreations` is
the only request-to-terminal record, so reconciliation looks up the
reservation per terminal and uses the same one for the destination
pane and the adoption. Reconciliation is suspended while a device
create is in flight, so a layout the host pushes before the receipt
returns waits until the reservation is bound.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that a growing grid replacement sheds droppable events first

A replacement Mac grid that outgrows the byte budget closes the
connection today even when queued terminal bytes could be shed to make
room. Normal admission sheds droppable events before it overflows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Admit a replacement grid through normal admission

A queued Mac grid is superseded by the next one for the same terminal,
so drop the old entry and admit the new frame the same way as any other
grid frame. Growing replacements now shed droppable events before an
overflow closes the connection.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that a queue overflow closes during lane negotiation

A drain that finishes a send while an independent-lane probe is parked
returns before it consumes a pending overflow, so the connection stays
open until the probe resolves.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Close an overflowed connection before yielding to lane negotiation

The drain now consumes a pending overflow before it checks for lane
negotiation, so a parked probe cannot delay the close.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Make the unbound device reservation tests deterministic

With one unbound reservation, a terminal no request created must not
take its pane. The adoption guard that protects the pane's queued input
is now tested directly on Workspace, since the placement test provider
never touches pending creations or the input relay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that a subscription negotiated across an overflow close does not leak

A queue overflow closes the connection while a second stream's lane
negotiation is parked. Close releases the connection's subscriptions, so
the subscribe that resumes after it must not register a new topic count
that nothing will ever release.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Drop a subscription whose connection closed during lane negotiation

The subscribe handler awaits the independent lane probe before it
registers the stream. A queue overflow can close the connection while
that probe is parked; close releases every subscription it knows about,
so a registration that lands afterwards leaks a process-wide topic count
and keeps the host emitting for a client that is gone. Re-check after the
negotiation and fail the request instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Materialize device terminals in the layout tests the way the device provider does

The split fixture created terminals through newTerminalSurface, which routes
to the machine instead of the viewer once the pane's selected tab is
Cloud-owned. A terminal bound to a reservation now takes the reserved pane,
and any other terminal gets a new manual-mirror pane at its destination, as
materializeManualMirrorTerminal does in production.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that a failed device reservation does not block layout updates

A failed create keeps its reservation and pane for Reconnect. Layout
reconciliation admits that panel through cloudPendingCreations, so a later
terminal still projects into its own pane and the reserved pane stays put.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that a pending or failed device reservation does not block layout sync

A reserved pane has no terminal on the owning Mac yet. The workspace must
still apply the owner's arrangement around it, send local gestures without
it, and close a mirrored pane's terminal on the owner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep a pending or failed device reservation from blocking layout sync

A reserved pane has no terminal on the owning Mac until its create binds.
Applying the owner's layout now grafts reserved panes back where this Mac
showed them, local gestures are sent without them, and closing a mirrored
pane no longer mistakes a reservation for an unrelated local pane.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that closing a connection ends an event drain parked in a write

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep the event drain's task handle so closing the connection cancels it

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Place stranded reserved panels in local order and bind every lent reservation in layout tests

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that a reserved pane keeps input typed before its first attach sticks

A reserved Cloud pane's first attach can fail while the Mac is still
reachable-but-not-ready. Input typed during that attach transition belongs
to the same remote surface and must be delivered in order once an attach
sticks. Input typed after an attached Mac disconnects is still dropped, and
stopping the mirror session (owner change) must discard queued bytes so a
replacement never inherits them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep a reserved pane's early input until an attach sticks

The mirror session discarded the adopted relay's queue on every transition
to .detached, including a failed first attach, so input typed while the
reserved pane was still attaching to its remote surface was dropped. The
relay now keeps that input for the same remote surface and delivers it in
order when an attach first sticks. After the handoff the device router
still drops input typed while detached, and stopping the session discards
anything held so a replacement owner never inherits it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that a lane that stays backlogged keeps its order storage bounded

A lane whose drain keeps pace without ever emptying leaves every consumed
ID in front of the order's head, where compaction never looks. After
10,000 enqueue/dequeue pairs the lane holds 10,045 IDs for 10 queued
events.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Count consumed IDs when compacting a lane's event order

Compaction compared only the IDs after the head against the live count,
so a lane whose drain kept pace without emptying never freed the IDs it
had consumed. Counting the whole array rebuilds the order once dead or
consumed IDs outnumber the live ones, which keeps the storage bounded and
every operation amortized O(1).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that a reserved pane drops early input when its Mac link drops

Input typed into a reserved pane is held across replay failures while the
link to the Mac stays up. Once the link drops before an attach sticks, the
Mac may come back with a new shell under the same surface ID, so the held
input must be dropped instead of replayed into it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Drop a reserved pane's held input when its Mac link drops

A reserved pane holds what was typed until an attach sticks, so a replay
that fails on a live link loses nothing. When the link itself drops first,
a restarted Mac can restore a terminal under the same surface ID with a new
shell, so the held input is discarded instead of replayed into it. The next
attach that sticks resumes forwarding.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Wait on the connection's close signal in the lane overflow test

The test polled the transport's close count for a fixed number of yields,
which can finish before a loaded runner's drain closes the connection.
It now awaits the connection's onClose, which runs after the transport
closes, under a one-minute test limit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that a rejected grid replacement keeps the queued grid

A replacement Mac grid that cannot fit overflows the connection. Until the
close runs, the queue should still hold the last admitted grid rather than
neither snapshot.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep a queued Mac grid until its replacement is admitted

The replacement used to drop the queued grid before checking whether the
new frame fit, so an overflowing replacement left neither snapshot queued
while the connection closed. The old grid's room now counts toward the
replacement, and the old entry leaves only once the new frame is admitted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Expect the beta nightly floor in the team What's New copy test

#12389 gave iOS 1.0.4 beta builds the 0.64.22 nightly floor, and
MobileMacCompatPolicyTests asserts it, but this copy test still expected no
nightly version. It only runs when the iOS simulator lane is routed, so it
failed on this branch's first full simulator run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Own event drains in the connection actor

Every drain now starts through one actor-isolated method that checks
isClosed and records the task handle in the same turn, so close() can
cancel every drain it admitted and a claim that arrives after close()
is released instead of started. This replaces the unfair lock that
guarded the handles from the nonisolated fan-out path; the fan-out
hops onto the actor once per claimed drain, not once per event.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Move reserved-panel layout grafting into CmuxCore

Removing and restoring reserved panels in a Mac device workspace layout
is pure tree logic, so it now lives on DeviceWorkspaceLayoutNode in
CmuxCore with package tests. Grafting indexes both trees once and wraps
a restored split around the lowest common ancestor of the panes it
divided, which is linear in the layout size instead of re-searching
the tree per restored panel. A randomized differential test checks it
against the previous per-panel insertion.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* State the unique panel ID precondition on layout grafting

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Read event order bookkeeping from the test target

The bounded-order test read a production accessor that existed only for
it. The arrival orders are now internal-read, and the test sums them
itself through @testable import. The large graft test no longer asserts
a wall-clock bound, which could fail on a loaded runner; it checks the
restored layout at 40,000 panels.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Drop main actor isolation from the grid publisher value

The publisher is a plain value that its owner mutates in place; its
closures run synchronously in the caller's context. Nothing about it
needs the main actor.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Index bound reservations once per device layout pass

Reconcile looked up each missing terminal's reservation by scanning every
pending creation, and filtered stale projections with a linear contains
over the wanted list. Build a key index of bound reservations and a set of
wanted terminals once per pass instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that a retried device create keeps the pane the layout already mirrored

A device split whose first create receipt is lost leaves its reservation
unbound, so the owner's layout mirrors the new terminal in a pane of its own.
Reconnect replays the create and gets the same terminal back. The test expects
that terminal to keep its single pane and later layouts to keep applying.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Reuse the mirrored pane when a device create returns a terminal already shown

A device layout can project the new terminal before its receipt binds the
reservation, after a lost receipt or when the layout event beats the create
response. The create's projection then made a second pane for that terminal
and every later layout for the workspace failed with an unmapped surface.
The reserved pane now gives way to the existing projection instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that a reserved pane giving way to its mirror hands over focus

When a retried device create finds its terminal already mirrored in another
pane, the reserved pane closes. If the user was in that pane (they pressed
Reconnect there), focus should land on the pane showing the new terminal
rather than on whichever neighbor the split tree picks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Hand focus to the mirrored pane when a reserved pane gives way

A device create that finds its terminal already mirrored closes the reserved
pane. If that pane had focus, the split tree picked a neighbor, which could
be an unrelated terminal. Focus now moves to the pane showing the terminal
the user asked for.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Reuse a mirrored projection only while its pane is still open

The catalog matches projections by resource, workspace and tab without
checking the panel. A create that found a projection whose pane had already
closed would finish on a pane that no longer exists and close the reserved
one. It now projects into the reserved pane as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that a device link delivers grid events to its surface's mirror

Terminal envelopes now reach their mirror sessions through one
DeviceLinkTerminalEvents.receive(_:) call, so the link has no second list
of terminal topics to keep in step with the decoder. The new test feeds
terminal.updated and device.terminal.grid envelopes through that path and
checks that each topic is subscribed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Find every restored split's layout anchor in one pass

Each restored split walked the owner tree from both of its terminals to
find their lowest common ancestor, so a chain-shaped layout with many
reserved splits cost the tree depth once per split. Grafting 32,000 nested
reserved splits around a 32,000-pane chain took 27.5 s.

A wrap only inserts a split above a target node, beside a branch with no
target panel, so the anchor from the unwrapped tree stays correct after
every wrap. Tarjan's offline algorithm finds all anchors in one pass over
the owner tree before grafting. The same layout now grafts in 0.25 s, and
8,000 splits in 0.06 s instead of 1.7 s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test grid delivery through the device link's own event routing

The previous test fed the envelope to the terminal fan-out directly, so a
topic case added ahead of the default branch in DeviceLink.handle could
swallow terminal.updated or device.terminal.grid without failing it. The
test now builds a DeviceLink and hands the envelope to handle, the method
the event consumer calls for every host event.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Remove catalog keys the main merge duplicated

The merge of main 83270f1 ran git's line merge on Localizable.xcstrings
because this clone had no xcstrings merge driver registered. The branch had
moved the three cloud.link.sshPreflight entries, so the line merge kept both
copies. This rebuilds the catalog with scripts/merge-xcstrings.py from main's
text: it matches main except for devices.link.error.notDiscoverable.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep main's key order in the string catalog after the merge

The merge's key-level union kept every string but moved keys out of
main's order, a 1,451-line diff. The catalog is rebuilt from main's text
with merge-xcstrings.py, so it differs from main only by the branch's
devices.link.error.notDiscoverable key.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that a remote replay resumes a hibernated agent terminal

Another Mac or the phone attaches to a terminal through
mobile.terminal.replay. When Agent Hibernation had torn the terminal's
runtime down, the replay came back empty and no output followed, so the
viewer showed a blank pane with no disconnect overlay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Resume a hibernated agent when a remote viewer attaches

A remote Mac or the phone attaching to a terminal is visiting it, the
same as selecting its tab on this Mac. mobile.terminal.replay now wakes
a hibernated agent before building the replay. Before, the replay of a
torn-down runtime was empty, no output followed, and the viewer showed
a blank pane with no disconnect overlay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Wake a hibernated agent only for a valid, bound replay

A replay with an invalid viewport report no longer wakes the agent
before it's rejected. Like explicit input, the resume goes through the
panel only when the resolved surface is still the panel's own, so a
respawn's outgoing panel can't be resumed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Name Settings › Devices in the not-discoverable error

Main's #14772 gave Devices its own Settings section and moved the
"Make this Mac discoverable" switch there, so the error's path to
Settings › Computers no longer matched a section. The English text and
all 20 translations now name the Devices section with the same words
main uses for its other Devices paths.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Six pages under /docs/cloud (overview, machines, workspaces and agents,
files and networking, CLI reference, security and troubleshooting),
wired into docs nav, sitemap, agent page index, docs search aliases and
the audited SEO matrix. Copy is checked against the CLI usage text and
the enforced plan limits; translated into all 20 site locales.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* zsh integration: use zsh/zselect for poll-loop sleeps (no fork)

The git-HEAD and PR-status watchers run one background loop per pane and
fork /bin/sleep once per second indefinitely. On a busy machine with many
panes this is a measurable, continuous source of process churn — which
also multiplies under host security/EDR agents that inspect every exec.

Replace the in-shell `sleep` calls with a `_cmux_sleep_cs` helper backed by
the `zsh/zselect` builtin (its `-t` timeout sleeps without forking), falling
back to /bin/sleep when the module is unavailable. This mirrors the existing
zsh/net/unix and zsh/parameter module-over-fork optimizations already in this
file.

Note: zselect with only -t and no fds returns status 1 on timeout, so the
helper calls it then returns 0 explicitly rather than falling through to the
/bin/sleep fallback (which would otherwise double the wait).

Verified: `zsh -n` passes; zselect -t 100 sleeps ~1.00s fork-free in a
backgrounded subshell with redirected std streams; fallback path sleeps
correctly when the module is absent.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: cover fork-free zsh watcher sleeps

* test: stabilize zsh watcher sleep coverage

* fix: normalize zselect timeout under strict zsh options

* test: verify zsh watcher teardown

* test: cover zsh sleep fallback watcher routing

* test: isolate zsh integration environment

* test: exercise watcher teardown with zsh job control

* test: synchronize zsh watcher readiness

* ci: retrigger PR validation

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Austin Wang <austinwang115@gmail.com>
…lback on a new line (#15152)

* Keep cmux's own Ghostty keys out of the config error card

Ghostty reports sidebar-font-size, surface-tab-bar-font-size,
sidebar-background and sidebar-tint-opacity as unknown fields because
cmux reads them itself. The card now drops diagnostics for those keys,
using the same key set the parser switches on, and shows nothing when
they are the only diagnostics.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* End restored scrollback on a fresh line

Captured scrollback usually stops at the old prompt with no trailing
newline, so the new shell's first prompt started mid-line and zsh
printed its highlighted % marker. Replay now appends CRLF when the text
does not already end on a new line (ignoring trailing CSI sequences).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Drop the extra newline after replayed scrollback in the disconnect placeholder

Replay now ends on a fresh line itself.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Basics, Workspaces, Agents (with the agent integrations), Automation,
Remote Access, cmux Cloud (now with Base), and Customize and Admin.
URLs are unchanged; only nav grouping and pager order change. Section
labels are translated for all 20 locales, and the unused
agentIntegrations label is removed.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… token is free (#15129)

* ci: run changed suites inside an owned compile admission when its gui token is free

A changed-suites run on a self-hosted Mac compiled in admission, then handed
the suites to a separate `app-host unit tests` worker, because admission held
no gui token. Over 89 first-attempt PR runs, 48 had exactly that one worker,
and the time from admission finishing to that worker starting its tests was
p50 544 s, p75 875 s, p90 1374 s (queueing for a gui runner, fetching the
product, restoring it). Running the same scripts in admission costs about
30 to 40 s (restore from the local archive, enumerate).

Admission now takes its Mac's gui token with take-gui, as test-e2e.yml's
build already does, and runs the suites itself when it gets one. When
take-gui gives way or times out, admission reports unit_tested=false and the
worker runs the suites exactly as before; macOS status then requires it.
Blacksmith runners have no helper and keep running the suites in admission.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: leave the suites to the worker on take-gui exit 2 or CI_PR_POOL_OWNED_GUI=0

Compile admission never gets the gui token at job start, so a helper that
cannot take it (exit 2) must not run the suites untokened. The owned-GUI
switch keeps them off the minis again, and the wait is bounded at the
hook's 240 s gui wait. A test runs the step against a fake helper.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: keep the gui-token step out of the product key; model unit_tested in the routing test

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* test(cmux-tui): replays must restore the OSC title

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cmux-tui): carry the OSC title through VT replays

Ghostty's VT formatter emits OSC 7 but never OSC 0/2, so every terminal
rebuilt from a replay (host resize, resync, daemon reattach, frontend
attach) came back untitled and the surface title was overwritten with an
empty string. Append an OSC 2 suffix beside the mouse-format suffix,
reserved in the replay budget and preflight.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…15176)

A hard link into RUNNER_TEMP gave the cached object a second link for as
long as the job held it. The LAN product server only serves single-link
objects, so a product was unservable to sibling consumers exactly while one
was using it, and the job could write through the link into the cache.
clonefile(2) costs the same on APFS and keeps the object single-linked;
elsewhere it still falls back to a copy.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: reproduce live Codex restore lease hang (#15111)

* fix: reject contended agent restores without queuing (#15111)

* test: cover slow restore retargeting and claim cleanup

* fix: bound all restore conflict presentation requests

* test: bound restore contention across repeated pane moves

* test: handle expected restore client disconnects

* fix: preserve bounded cleanup after late restore admission

* test: await kernel lease release after watcher exit
* ci: bound the release-gate launcher

* ios: report active v2 peer transport path
…15159)

* test: expect read_text to name why a terminal is not running

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* read_text: explain why a terminal is not running and how to wake it

surface_unavailable now carries data.reason (hibernated, awaiting_restore,
starting, closing) and a message per cause. A hibernated terminal also gets
data.wake_command, the focus-panel command that shows and wakes it. A restore
that waits for admission after cmux reopens now waits like a starting
terminal, because it starts by itself once the agent session is known.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* read_screen: share the not-running explanation with read_text; translate all catalog locales

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…it with the ping (#14872)

Three changes to the work Resources/bin/cmux-claude-wrapper does before
it execs claude:

- Generated hook settings that are byte-identical to what
  `cmux hooks claude inject-settings` prints skip the Node validation.
  Any other bytes (older/newer CLI, a different cmux on PATH, garbled
  output) still go through Node unchanged.
- The managed-policy check reads the com.anthropic.claudecode domain
  once per launch instead of four times (two keys, asked twice), with
  the same decisions as the per-key reads.
- The settings generator runs in the background while the socket ping
  decides whether this launch gets hooks, and every passthrough
  discards it.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add failing tests: cached update relaunch save drops tmux reattach bindings

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep tmux reattach bindings when a session save has no process scan

The update relaunch save and the quit watchdog fallback after a timed-out
fresh scan use ProcessDetectedResumeIndexes.cached, whose binding index is
unavailable. The snapshot projection treated that as a clean empty scan and
dropped every process-detected binding (tmux, SSH), and the Dock also
deleted its stored copy. Treat an unavailable index as missing evidence and
persist the binding the last successful scan stored.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* ci: post a one-click dogfood link on app pull requests

Each same-repo pull request that changes the app or CLI gets one sticky
comment with a link to a tagged dev build of its exact head. The build
controller treats the job's success as the build request and verifies the
pull request before building. The job is not required and not part of
ci-status.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: quote the dogfood job name and keep comment failures non-fatal

The unquoted ` #` started a YAML comment, so the job reported as
"Dogfood build" without its number. Label events no longer re-request a
build, and a failed comment no longer turns the run red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…download earlier (#15175)

cp -cR clones a tree file by file (about 12 s of kernel time per 50,000
files); one clonefile(2) on the directory makes the same tree in about 1 s.
Keep, owned adopt, the owned package clone, local seed clones and the
canonical source copy now try it first and fall back to the old copy.

The DerivedData seed download needs only the fingerprint, so it now starts
before the GhosttyKit, Swift package and manifest cache restores instead of
after them.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
teamleaderleo and others added 21 commits September 28, 2026 13:00
…again (#15061)

* tests: find the Settings window by identifier, not title

Most SettingsUITestCase tests fail at "Settings window did not open" on
CI, while the tests that look the window up by its identifier
(cmux.settings) and wait 8 s pass. Use the identifier everywhere and give
the first open 10 s, since Settings mounts its sections progressively.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* tests: activate the push-forwarding Settings test like its siblings

It launched without activation, so with the window now found the scroll
into Mobile had no hit point. Use launchAndActivate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* tests: open Settings after activation in the push-forwarding test

It opened Settings at launch, and activating the app then put the main
window over it, so the sidebar scroll view had no hit point. Launch and
open Settings with ⌘, like the other Settings tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: document shared Settings UI test setup

* test: resolve nested Settings toggle controls

* test: read selected Settings toggles reliably

* tests: keep an identifier-bearing element as the toggle lookup's last resort

The nested switch probe made toggle() fail for Settings rows that expose
neither a switch nor a checkbox. Fall back to the first element carrying
the identifier, and use firstMatch so duplicate ids cannot raise.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* tests: launch Settings App tests in a fresh home; read toggle value first

The Inherit Working Directory and all-surfaces toggle tests failed on the
fleet minis because an earlier run's flipped values survived resetDefaults,
so the toggles started in the wrong state. SettingsAppBehaviorUITests now
launches the app with its own HOME, CFFIXED_USER_HOME and XDG_CONFIG_HOME,
so every test starts from factory defaults and leaks nothing.

isOn reads the accessibility value and uses isSelected only when a control
reports none. The localized search test selects and retypes through the
app, since typing can replace the search field's accessibility element.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* tests: use a toggle's own identifier-bearing element when it reports a value

With a fresh home the Inherit Working Directory and all-surfaces toggles
still read the wrong state. They passed before the lookup started
preferring a switch or checkbox nested under the identifier-bearing
element, so that element is used again whenever it reports a value; the
nested control is only the fallback for a valueless container (the
push-forwarding toggle). Start-state failures now print the resolved
control's type, value and selection.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* tests: flip the inherit and all-surfaces toggles from their current state

A fresh HOME and CFFIXED_USER_HOME did not help. The resolved checkboxes
still read inherit=0 and all-surfaces=1 at launch, the opposite of the
catalog defaults and exactly what these tests leave behind. On the fleet
minis the value survives resetDefaults. The two tests now start from
whatever state the toggle reports, check that each click flips it while
the subtitle stays, and flip it back so the machine keeps its setting.
The home isolation is removed. The toggle lookup uses the element that
carries the identifier when it reports a value.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* tests: flip Minimal Mode from its current state; pin standard layout for Help menu tests

A fleet mini kept Minimal Mode on from an earlier run, and resetDefaults
did not clear it. Minimal Mode hides the footer's Help button, so all
three SidebarHelpMenuUITests failed at 'sidebar help button', and the
Minimal Mode test failed at its start-off assertion. The Help menu tests
now launch with -workspacePresentationMode standard, as
testCommandPaletteCanEnableAndDisableMinimalMode already does. The
Minimal Mode test flips from whatever state it starts in and flips back.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…5359)

* test(cloud): cover the reads that find a Cloud machine gone

A status read and an access preflight both move the row to `destroyed`
themselves when the provider no longer has the machine. Once either does,
`destroyVm` can never see the row again (its lookup skips destroyed rows)
and the provider-status cron skips it too, so the model-plane revoke and
the `vm.destroyed` usage event that the cron performs for the same
transition have to happen at these write sites as well.

Fails today: both retire the row and record nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): finish the destroy work where a machine is first found gone

Three places retire a Cloud machine's row after the provider says it no
longer has the machine: the status read, an access operation's resume
preflight, and the stats read. Each wrote `status = destroyed` and stopped
there, while the reconcile cron doing the same transition also revoked the
machine's model-plane tokens and recorded a `vm.destroyed` ledger event.

That difference was permanent, not a race to lose. `destroyed` is terminal:
`findUserVm` hides such a row from every destroy request and
`reconciliationCandidates` drops it from the cron, so whichever of the three
got there first left a machine that never appears as destroyed in the ledger
and never has its route tokens marked revoked, with nothing able to finish
the job afterwards.

All three now go through one `applyObservedProviderStatus` that performs the
write and, when the write lands on `destroyed`, the same revoke and ledger
event as the cron. The status route hands `getVm` the model-plane revoker it
already builds for delete. The new `provider_status_*` destroy reasons join
the analytics allowlist so the event says which read noticed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): derive the observed status inside the shared destroy path

Review found that the previous commit took the row's new status from each
caller. That let the access preflight and the stats read hardcode
"destroyed" for a provider 404, including for a machine with a persistent
home volume, which observedDbStatus maps to "paused" because the compute
is gone and the machine is not. Those rows were terminalized and billed a
vm.destroyed that never happened, and nothing revisits a terminal row to
take either back.

The status is now derived inside applyObservedProviderStatus, so every
entrypoint agrees about what a 404 means. reopenBaseIfProviderDeleted is
the one caller that must override it, and passes forceStatus with the
reason: its row is a Base's active generation, and leaving it paused would
hand the same dead provider id back on every later open.

Also threads the model-plane revoker through getVmStats from its route,
covers the stats entrypoint including the home-volume case, and fixes the
two test-file regressions the previous commit shipped.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* vm: say what a paused 404 row means for credentials

The access preflight carried a comment claiming a revoke was pointless
there because the credentials were already inert. That is false for the
case this branch introduces: authenticateRouteToken and
authenticateVmAuthorization both accept `paused`, so route tokens on a
volume-backed machine stay valid after this write, for the rest of their
30-day lifetime.

Keep the behavior, which is what getVm and the reconcile cron already do
for the same observation, and replace the comment with what is true. Also
drop the stale "next fleet refresh drops it" line, correct "seven call
sites" to eight, stop asserting a resurrection path that is not
implemented, and type usageEventSource as VmDestroySource so an unknown
source cannot silently degrade in PostHog.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* ci(seed): keep the trusted seed on the Mac before the R2 upload

The LAN seed archive reads kept seeds from the trusted seed Mac within a
minute of the keep. Keep ran after "Save seed", so every LAN seed waited on
the R2 upload first (p50 160 s on 2026-09-28). Keep only clones the tree, so
the kept copy is still exactly the seed the upload tars.

cmuxterm-hq#658.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: say why the trusted seed job keeps before it saves

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ry app PR (#15418)

* PR media: run for every app pull request, not only dev-build ones

#15380 made the dogfood build job opt-in, and the media gate waited for
its success, so only labelled pull requests got media. The gate now
waits for the dogfood job only when it runs (it rewrites the sticky
comment) and otherwise decides on the app build alone; publish posts the
sticky comment itself and updates a media-only comment for a new head. A
pull request that changes nothing a tour shows gets no media section.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* PR media: adopt only, start when CI completes

Tours never compile on their own any more: each adopts the build CI made
for the pull request on the pool that compiled it, or main's build of the
same inputs when CI reused it (dispatch-focused-test.py --adopt-main, which
lets test-e2e.yml adopt main's product and fail rather than compile). A
build the UI test Macs cannot load gets a skip note; only a manual
allow_compile dispatch compiles.

The workflow starts when the CI run completes instead of when it is
requested, so the planner no longer holds a runner while CI builds. A
dogfood comment naming an older head no longer holds media back.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* PR media: tour main's build on the merge CI tested; address review

- A tour of main's build dispatches the merge CI tested, so test-e2e.yml
  looks main's product up by the inputs CI matched, not the head's.
- A manual dispatch while CI still runs leaves media to the completed run.
- Skip notes say how to retry now that nothing retries by itself.
- Pin include-hidden-files: false on the media upload.
- Drop stale compile comments; test the --adopt-main parser guard.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: cover the Cloud toolbar's account of a failed machine-list read

With cached machines still on screen, the toolbar row renders one warning
triangle and one "Machine list unavailable" line for all three failures and
offers nothing to click. The notice and the empty state already route
.unreachable to Retry, .sessionRejected to a fresh sign-in and .requiresPro
to an upgrade, from the same MachineListStatusPresentation.

This commit adds the failing coverage and the `perform` parameter the row
needs to reach the handler. The body still hardcodes the triangle and the
generic string, so the new tests are red.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: let the Cloud toolbar name the failure it has and offer its fix

MachinesListStatusToolbarRow now reads MachineListStatusPresentation the way
the notice and the empty state already do, instead of hardcoding one warning
triangle and one "Machine list unavailable" line. A rejected session shows the
account glyph and a Sign In link; a lapsed plan shows the Pro glyph and an
Upgrade link; an unreachable read keeps Retry. Both go through the same
performListStatusAction the other two surfaces use.

The presentation gains `staleTitle`, the one-line form for the toolbar, so the
shorter copy still lives beside the full copy it has to agree with, and Action
gains `shortTitle` because the toolbar has one line beside the status text.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: keep the Cloud toolbar's failure affordances off the non-failure states, and make the tests discriminate

Review follow-up.

The row applied `.help(failure ?? "")` and `.cloudErrorCopyMenu(failure)`
unconditionally, so waiting and reconnecting got an empty tooltip and,
worse, an empty `.contextMenu {}`, which on macOS suppresses whatever menu
that region would otherwise inherit. Base applied neither to those
states. Both now hang off the failure branch. The action button also
takes `.fixedSize()`: in a narrow sidebar, truncating the status line is
survivable, losing the only affordance that fixes the failure is not.

Three test problems, all of which let a regression through:

- Nothing covered `failure = presentation.isFailure ? error : nil`, the
  one genuinely new branch. Simplifying it to `let failure = error` kept
  every test green while offline gained an orange dismissable chip. The
  offline case now asserts no `CloudBannerDismissButton` for waiting or
  reconnecting, and one for each failure.
- "The three failures do not read the same line" compared the three
  rendered strings pairwise, but each row already carries a distinct glyph
  and a distinct button, so three identical sentences would have passed.
  Each row is now matched against its own `staleTitle`, the three stale
  lines are checked to be distinct, and the panel's paragraph is checked
  not to leak into the toolbar's single line. That subsumes the separate
  stale-qualifier test, which is gone.
- Assertions used the English literals "last known" and "Offline", which
  a copy edit or a non-`en` host would have reddened for no behavioral
  reason. They read the presentation and the catalog now.

Also removed `.environment(\.accessibilityEnabled, true)` from the test
host and the comment claiming it switched SwiftUI's accessibility output
on. It does not: the key is a read-only signal for app code, and it is
deprecated. The window is what populates the tree, which the comment now
says.

## Changelog

Fixed: the Cloud sidebar toolbar no longer suppresses the header's context menu while waiting for the network or reconnecting.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: restore the accessibility environment the toolbar suite renders through

Run 36401958401 failed every test in this suite with nil elements and
empty text. The previous commit removed `.environment(\.accessibilityEnabled,
true)` on the claim that it did nothing and a window was the real
requirement. The window is not sufficient: the same assertions passed in
run 36397834894 with the modifier and no window, and fail with a window
and no modifier. In-process there is no assistive client to switch
SwiftUI's accessibility output on, so the hierarchy has to ask for it.

The window is kept, since querying a detached hosting view is worth
avoiding on its own, and the comment now records which of the two is
load-bearing so it is not removed again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…re (#15291)

* Test that only background upkeep waits out the Cloud link backoff

Red: CloudMachineLinkManager refuses every connect for 15 s after a
failure, including a person's open, and there is no upkeep mark.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Let a Cloud open dial through the link backoff; only upkeep waits

After any failed link, CloudMachineLinkManager refused every connect to
that machine for 15 s with the old error, including a person's open,
which then showed "The Cloud connection did not complete" without
trying. The backoff now applies only under CloudLinkUpkeep, which the
sidebar's periodic refresh sets. Opens, terminal creation, retries and
socket commands dial. Also runs the CmuxCloud package tests in the
package test lane, which never included them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep the upkeep mark on the link manager

The package conventions lint rejects an all-static namespace enum, so
the task-local lives on CloudMachineLinkManager as isBackgroundUpkeep.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Run the backoff tests in the app-host suite

On a pull request the package lane routes from the base branch's list,
so CmuxCloud's tests never ran here. The tests move to cmuxTests, which
the changed-suites job runs, and the lane list stays as on main.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…15412)

* ci: ask for the fuzz regression replays on sidebar, split and window changes

choose_ci_suite.py adds cmuxUITests/FuzzRegressions to ui_selectors when a
same-repository pull request touches a path the UI fuzzer's checked-in
repros exercise (ui_tests_dispatch.FUZZ_REGRESSION_PATHS), so the ui-tests
job replays them in the UI test lane against the app it already adopts.
The replays count toward one focused run's selector limit and never close
a cmuxUITests/ coverage gap.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* ci: keep a UI class run whole when the replay does not fit; harden the replay

The fuzz replay gives way when the changed classes already fill one
focused run, instead of turning them into a coverage gap. The window
resize step waits as long as the app does (30 s), and a repro that stopped
before its last step fails instead of passing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fuzz: fail a replay on a pointer-error step; say which pull requests get the replay

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Test that an equal-cursor conflict cannot wedge a Cloud graph

Red: a full snapshot that disagrees with the applied graph at the same
cursor is refused forever, and there is no equalCursorConflict state.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Break an equal-cursor conflict wedge on the second full read

A full snapshot that disagreed with the delta-built graph at the same
cursor was refused forever, so the machine's graph never became current
again. The first conflict from a current full refresh now keeps the
graph and schedules the bounded recovery read; a second conflict at the
same cursor adopts the daemon's snapshot and leaves a breadcrumb.
Event-feed snapshots and stale reads never adopt.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Arm-only recovery reads; clear the armed conflict on delta and suspend

Only the install that arms an equal-cursor conflict schedules the
recovery read, so stale or rename-fenced reads do not spend the budget.
A delta that advances the cursor and a feature-flag suspend both clear
the armed conflict.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep the test catalog alive for the provider's unowned reference

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Assert the conflict is armed before clearing; cover equal-content clearing

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(panes): make split admission ratio-aware

* fix: centralize split-space enforcement

* fix(panes): trace prospective mixed splits

* fix: preflight splits before panel startup

* fix: preserve browser fallback before split preflight
Tours could click and hover but not drag, so a resizer or any other
drag handle could not be exercised from a tour. dragAt presses at one
point in the window and drags to another, both in the same 0 to 1
window space the existing clickAt and hoverAt steps use.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add failing test: cursorless snapshots freeze the Cloud graph

Two snapshots from a daemon that sends no cursor compare nil == nil, take
the equal-cursor branch, and the second one is refused whenever its
content changed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Treat two missing cursors as unordered, not equal

A daemon that sends no cursor made every later snapshot take the
equal-cursor branch, so the first content change froze that machine's
graph for the rest of the session. Require a present cursor there; the
snapshot decision already installs a cursorless snapshot over a
cursorless graph and refuses it over a journaled one.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep reads ordered by install version when neither side has a cursor

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* ci: scope contributor web complexity to web changes

* test: update workflow filter parity for scoped candidate

* ci: preserve candidate workflow trust boundary
* Add btop-style agent activity custom sidebar example

btop-agents.js lists every workspace with a braille sparkline of recent
agent activity, a state glyph, a small progress meter, unread count and PR
number, under a figlet header with a graph of busy workspaces. History is
sampled per clock tick into per-workspace buckets inside the sidebar.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep btop sidebar graphs readable on light sidebars

System yellow braille dots were faint on a white sidebar; the heat ramps
now go green to orange to red.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(example): bound btop history across clock changes

* fix(example): cap btop workspace modeling

* fix(example): retain btop history across filters

* fix(example): bound btop history with LRU eviction

* fix(example): cache bounded btop workspace selection

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* dogfood: make modifier-clicks-tour click on its own URL

This tour has been green since it landed and has never exercised a modifier
click. Measured against its own frames from run 36435029937, where the window
is 960x487, the row pitch is 9 px, the first output row is at y=33 and the
character width is 5.5 px:

  - the tour prints one line of output, holding one URL that ends at x=340
  - it clicks at (384, 49), (384, 68) and (384, 88)

Every click is past the end of the text and one to six rows below it, with
nothing underneath. clickAt reports ok whether or not anything is under the
point, so the tour passes on a window where nothing can happen. I built a
hypothesis about Command not being delivered on top of that green result and
filed an issue asserting it; both were wrong, and the tour is why.

Print 60 rows so a click has a large target, and click at x=0.01, y=0.40,
which lands inside the first URL of about row 18. Two panes with two different
URLs, so a frame identifies which click produced it and one opened browser
cannot be mistaken for the other.

Also drop the six hover shots. shot() calls window.screenshot(), or
XCUIScreen.main.screenshot() for a screen shot, and neither draws the mouse
cursor. cmux signals a link by turning the pointer into a pointing hand, so a
cmd-hover frame is pixel-identical to the plain frame before it. Those shots
could never show the thing they were named for, and they crowded the frames
that carry the proof out of pr_media's four key-shot slots.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* dogfood: restore modifier hover coverage

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* ci: register the nightly owned-Mac producer on the fork default branch

GitHub only dispatches workflows that exist on the default branch; the
content that runs comes from the dispatched ref.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: reject masked Swift Testing failures

* ci: leave the retired nightly mini lane removed

* ci: distinguish Swift Testing issue records

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…15419)

* Merge project.pbxproj by union of added entries instead of by line

No target in this project is filesystem-synchronized, so adding one source
file writes four pbxproj entries: a PBXBuildFile, a PBXFileReference, a
group child and a Sources phase member. Two branches that each add a
different file append to all four of the same regions and collide
positionally, even though the entries are disjoint. It is the most common
conflict in the repository and it is never a disagreement.

scripts/ci/catch_up_pr.py already resolves exactly this, with union_pbxproj.
That only ever helped the catch-up bot and merge-main.sh. It did not help a
local `git merge main`, a rebase, or a pull request from a fork, which
catch-up refuses by design and which therefore re-conflicts on this file
every time main moves.

Expose the same union as a git merge driver, registered the way the
.xcstrings one already is: an attribute in .gitattributes and a config line
in install-git-hooks.sh, which setup.sh runs. Existing clones pick it up by
re-running that installer.

The driver is conservative in the same way the union is. A hunk merges only
when both sides purely added distinct lines; anything else returns non-zero
and git writes normal conflict markers. Because the driver overwrites the
working-tree file, the union is also run through normalize-pbxproj.py on a
temporary copy first, and a result that would not open in Xcode is refused
rather than written.

catch_up_pr.py pins merge.pbxproj.driver=false alongside the xcstrings one.
Without that the bot would execute a merge driver script out of the tree it
is merging, which on a fork head is untrusted code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(ci): materialize rejected pbxproj conflicts

* fix(ci): restrict pbxproj unions to source entries

* fix(ci): validate pbxproj merge structure

* fix(ci): parse pbxproj merge context structurally

* fix(ci): parse pbxproj section comments structurally

* fix(ci): harden installed merge drivers

* fix(ci): pin the merge driver interpreter

* fix(ci): canonicalize merge driver trust roots

* fix pbxproj merge driver conflict fallback

* fix(ci): isolate and version merge drivers

* ci: register pbxproj merge regression test

* fix(ci): migrate legacy merge driver configs

* fix(ci): harden merge driver migration

* fix(ci): install immutable merge migration hooks

* fix(ci): verify merge installer inputs

* test(ci): update trusted hook fixture

* fix: keep project merge hooks trusted and precise

* fix: verify every installed hook input

* fix: always materialize complete merge conflicts

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
createMainWindow copies the size of the current main window, and earlier
tests in the app host register 320-point main windows. With the sidebar,
the new window's split container is then narrower than two minimum-width
columns, so split admission (#15392) correctly refuses the side-by-side
split these tests make. Size the window and its split container before
splitting instead of weakening admission.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Catch-up merge by scripts/ci/catch_up_pr.py (RFC #14631).
Merged by scripts/merge-main.sh: origin/main at ce5cb45.

Catch-up-previous-head: 6ce2e23
Catch-up-base: ce5cb45
The package-conventions lint rejects public enums that hold only static
members, and ControlWorkspaceRemoteLocalSocketPath was one. That failed
package-conventions-lint and the release-ios convention guard and every
job gated on them.

The resolver is now a struct that takes the controller's socket path at
init, and resolved(requested:) returns that path when the client asks
for forwarding. The caller injects currentSocketPathForRemoteRestore().
Behavior is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@cursor

cursor Bot commented Sep 28, 2026

Copy link
Copy Markdown

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@github-actions

Copy link
Copy Markdown
Contributor

Dogfood tours of cae5f924

sidebar-and-chrome-tour at cae5f924: not run

skipped: CI built this head on a runner pool whose products the UI test Macs cannot load, and media never compiles one; gh workflow run pr-media.yml -f pr=&lt;n&gt; -f allow_compile=true does

agent-activity-reorder at cae5f924: not run

skipped: CI built this head on a runner pool whose products the UI test Macs cannot load, and media never compiles one; gh workflow run pr-media.yml -f pr=&lt;n&gt; -f allow_compile=true does

Tours are picked by the paths globs in dogfood/scenarios/*.json; a Dogfood-tours: a, b line in the description picks them instead (none turns this off). Look at every frame before merging: a green tour only means no step failed.

The foreign-peer tests close the client socket without writing. When the
fake server thread accepts after that close, setsockopt(SO_NOSIGPIPE) on
the accepted socket fails with EINVAL, and the server's reply write kills
the test process with SIGPIPE. swift test for CmuxRemoteWorkspace died
with signal 13 on every full parallel run.

Set the option on the listening socket instead; accepted sockets inherit
it from the moment they are created.

Red: swift test --package-path Packages/macOS/CmuxRemoteWorkspace exited
with signal 13 on 4 of 4 runs. Green: 158 tests in 25 suites passed on
6 of 6 runs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
austinywang added a commit that referenced this pull request Sep 28, 2026
…-path

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

# Conflicts:
#	Packages/macOS/CmuxRemoteWorkspace/Tests/CmuxRemoteWorkspaceTests/RemoteCLIRelayServerTests.swift
#	Packages/macOS/CmuxRemoteWorkspace/Tests/CmuxRemoteWorkspaceTests/RemoteDaemonProxyTunnelCloudCLITests.swift
@austinywang

Copy link
Copy Markdown
Contributor Author

Closing as superseded by merged PR #15116, which consolidated this SSH/security/local-state hardening into main.

@austinywang austinywang closed this Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants