Skip to content

Make agent key hints clickable in terminal panes - #17765

Closed
azooz2003-bit wants to merge 166 commits into
feat/agent-action-buttonsfrom
feat/agent-key-hints
Closed

azooz2003-bit wants to merge 166 commits into
feat/agent-action-buttonsfrom
feat/agent-key-hints

Conversation

@azooz2003-bit

Copy link
Copy Markdown
Collaborator

Summary

Agent TUIs print key hints all over their output: ctrl+o to expand under a collapsed tool result, esc to interrupt in the status line, shift+tab to cycle beside the mode. With the new agentActions.keyHints setting on (off by default, Settings > Terminal > Clickable Agent Key Hints), clicking one of those hints in a terminal pane presses its keys. Hovering a hint underlines it, shows a pointing hand, and a tooltip such as "Click to press ⌃O (expand)".

How it works:

  • Detection. AgentKeyHintDetector (CmuxTerminalCore) reads the clicked viewport row on its own (no joining of soft-wrapped lines): one or more chords (ctrl+x, alt+x, shift+tab, esc, tab, enter, space, arrows, ⌃O-style glyphs, ctrl+x ctrl+k sequences), optionally again, then to/for and an action that ends at ·, |, parentheses, a two-space column gap or the end of the line. OpenCode's bare esc interrupt form matches only for its known actions. Bare key words (up to date, end to end) count only before a hint verb, and return is not a key word. Keys that act on whatever UI is live (Enter, Esc, Tab, arrows, bare letters, Home, End, page keys) count only in the agent's live region: the viewport is at the bottom and the row is no more than 8 rows above the cursor row, or below it, where the status line, input box, footer and permission prompts are drawn. Ctrl and Alt chords (ctrl+o to expand) count on any row, so prose such as "Scroll down to view the full log" stays inert.
  • Where. Only in panes whose journaled agent lifecycle includes Claude Code, Codex or OpenCode, and only while the setting is on. When several are journaled, a running or waiting agent wins over an idle one. Nothing runs on the typing path; the click is resolved on mouse release, and hover reads one row only when the pointer moves to a new cell. Scrolling, a resize, or a viewport change under a shown hint clears or re-resolves the underline.
  • Safety. Chords that quit, suspend or signal (ctrl+c, ctrl+d, ctrl+z, ctrl+\) are never clickable, and are never sent even if a keybindings file maps a hinted action to one. Any chord with Ctrl and c, d, z or \ counts, whatever Shift or Alt is added. A drag, a selection or a Shift/Option/Control click presses nothing. A click presses its hint once the system double-click interval passes without a second press, so a double or triple click selects text and presses nothing.
  • Click vs Command-click. If the agent has not captured the mouse, a plain click presses the hint. If it has, the click belongs to the agent and only a Command-click presses the hint. A Command-click on a hint takes over from the link and path fallback only when a hint covers the clicked cell; links, paths and SHAs elsewhere open as before.
  • Binding layers. For Claude Code, a hint printed with its default chord (expand, cycle mode, run in background, interrupt, amend) is remapped to the user's chord when ~/.claude/keybindings.json binds that action to a different one. The file is checked on click and cached by modification date; hover reuses a check for up to 5 seconds. Codex and OpenCode render their hints from their own keymaps, so their printed chord is sent. Then, if the pane's Ghostty keybindings would consume the chord (ghostty_surface_key_is_binding), the key's legacy encoding is sent instead (ctrl+o as 0x0f, shift+tab as ESC [ Z, arrows as ESC [ A...), so the terminal binding doesn't run; a key with no legacy encoding is skipped.

OS-level remaps (Karabiner, modifier keys in System Settings, hidutil) and a tooltip that names the physical key to press come in a follow-up.

Part of #15223. Stacked on #15234.

Testing

Ran locally:

  • swift test --package-path Packages/macOS/CmuxTerminalCore --filter 'AgentKeyHint|TerminalLegacyKeyEncoding': 29 tests in 6 suites passed (detector incl. live region and prose, signal chords with extra modifiers, chord resolver against keybindings files, mtime cache and hover max age, live region rule, deferred press timing, click policy, legacy encodings).
  • swift test --package-path Packages/macOS/CmuxTerminal --filter TerminalTextRegion: passed (a viewport row selects exactly that row; a blank row below short content reads blank; a soft-wrapped line reads one row at a time).
  • swift test --package-path Packages/macOS/CmuxSettings --filter SettingCatalogTests: passed, including the new default-off test.
  • python3 tests/test_cmux_schema_parity.py and python3 tests/test_cmux_settings_supported_paths.py: passed.
  • python3 scripts/verify-local.py: 15/15 checks passed. python3 scripts/localization_catalog.py check: 0 parity errors.

Not run locally: the app compile and cmuxTests. CI builds the app and runs the new cmuxTests/TerminalAgentKeyHintTests.swift, which checks that a hint click in an agent pane enqueues one key (two for ctrl+b ctrl+b), that nothing is clickable with the setting off, that bare keys are clickable only in the live region, that a running agent wins over a stale lifecycle, without an agent lifecycle, on prose or ctrl+c lines, and that a mouse-capturing agent needs a Command-click. The mouse release, hover underline, tooltip and Ghostty binding fallback are not yet exercised in a live build.

Localization: new Settings title and subtitle in the CmuxSettingsUI catalog, and the Settings search title and alias plus both tooltips in the app catalog, all in en, de, fr, ar, es, zh-Hant, zh-Hans, ko and ja.

Changelog

Added: Key hints that Claude Code, Codex and OpenCode print in a terminal, like "ctrl+o to expand", can be clicked to press those keys (Settings > Terminal > Clickable Agent Key Hints)

Demo Video

Pending dogfood.

Checklist

  • Behavior changes have added or updated tests, or Testing says why not
  • UI, settings, menu, schema, help-text or user-facing docs change: localization audited, and the result is stated above
  • User-facing docs updated if needed
  • Reviewed with a subagent before merge (cmux-review), and all bot and human review comments resolved

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.


Summary by cubic

Makes the key hints Claude Code, Codex, and OpenCode print in terminal panes (like ctrl+o to expand or esc to interrupt) clickable, behind a new off-by-default setting at Settings > Terminal > Clickable Agent Key Hints. Clicking a hint presses its keys; hovering underlines it and shows a pointing hand and a tooltip. The exact /goal resume prompt Codex prints is also clickable, sending the command and Enter through the same deferred hint path.

How clicks behave

  • Keys that quit, suspend, or signal (ctrl+c, ctrl+d, ctrl+z, ctrl+\) are never clickable or sent, even when a keybindings file maps an action to one, including with Shift or Alt added.
  • A plain click presses a hint once the double-click interval passes, so a double or triple click selects instead of pressing; a pending press is re-validated against the current runtime surface, generation, row, and authorization just before sending, and survives pointer exit. When the agent captures the mouse, only a Command-click presses the hint, and it wins over the link and path fallback for that cell.
  • Codex's /goal resume prompt clicks only in the live row region with an allowlisted command; no other shell text is clickable.
  • Bare-key hints (Enter, Esc, Tab, arrows, Home, End, page keys) count only in the agent's live region—viewport at the bottom of the scrollback, within 8 rows of the cursor. Ctrl and Alt chords count anywhere; "return" is never a key word.
  • Claude Code hints printed with their default chord are remapped to the user's chord from ~/.claude/keybindings.json, read on click and cached by modification date.
  • When the pane's Ghostty keybindings would consume the chord, the key's legacy terminal encoding is sent instead so the binding doesn't run; keys with no legacy encoding are skipped.
  • Hints resolve from the exact viewport row under the pointer, so a blank row can't map to content above it; hover re-reads per new cell and re-resolves on scroll or layout.
  • This branch carries a long stale merge of main; only the agent key hint changes belong to this PR.

Written for commit b231717. Summary will update on new commits.

Review in cubic


Migrated from #15305 after correcting the PR author identity. The head branch and commit history are preserved.

teamleaderleo and others added 30 commits September 28, 2026 05:12
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Behind the new agentActions.keyHints setting (off by default), a key hint
that Claude Code, Codex, or OpenCode prints in a terminal pane, such as
"ctrl+o to expand" or "esc to interrupt", presses those keys when clicked.
Hovering a hint underlines it, shows a pointing hand, and a tooltip names
the keys. When the agent has captured the mouse, the click belongs to the
agent and only a Command-click presses the hint; a Command-click on a hint
then wins over the link and path fallback for that cell.

What gets sent respects both binding layers. For Claude Code, a hint
printed with its default chord is remapped to the user's chord when
~/.claude/keybindings.json binds that action elsewhere (read on click,
cached by modification date). If the pane's Ghostty keybindings consume
the chord, the key's legacy terminal encoding is sent instead so the
binding does not run.

Part of #15223.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…/agent-key-hints

# Conflicts:
#	cmux.xcodeproj/project.pbxproj
…s_binding

AgentKeyHintDetector, AgentKeyHintChordResolver and AgentKeyHintClickPolicy
become values with their inputs as stored properties, and the legacy key
encoding becomes a String initializer, so the package namespace-type lint
passes. Package tests link the Ghostty runtime stubs, which now define
ghostty_surface_key_is_binding.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…turn

Ctrl+Shift+C, ctrl+alt+c and ctrl+shift+z still send 0x03 and 0x1a but
normalize to chords outside the blocked set, and "return" reads prose
such as "then return to continue the loop" as an Enter hint.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Read exactly the pointer's viewport row (TerminalTextRegion.viewportRow)
  instead of bottom-aligning the trimmed, unwrapped viewport text, which
  mapped a click on a blank row to a line of content above it.
- Hover reads that one row per new cell, reads the setting once per event,
  and reuses a keybindings file check for up to 5 seconds.
- Press a clicked hint once the double-click interval passes without a
  second press, so a double or triple click presses nothing.
- Bare keys (Enter, Esc, Tab, arrows, letters, Home, End, page keys) count
  only in the agent's live region: viewport at the bottom, row no more than
  8 rows above the cursor. Ctrl and Alt chords count anywhere. "return" is
  no longer a key word.
- Block any Ctrl chord on c, d, z or backslash whatever Shift or Alt is
  added, in the detector and the keybindings resolver.
- Clear the hover on scroll and layout, and re-resolve it when the viewport
  under a shown hint changes.
- Prefer a running or waiting agent over a stale idle lifecycle.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat: make Codex goal resume clickable

Add a narrow, live-row allowlist for the exact /goal resume prompt. A click sends the command and Enter through the existing deferred agent hint path.

* feat: show hover affordance for Codex action commands
…gation

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…e navigation

A pathless hub route puts Plan & billing and every team page beside the
Settings subnav (Account, Billing, Teams) while /dashboard/billing and
/dashboard/teams/* keep their public URLs. The sidebar keeps one Account
entry, Settings, current on every hub page. Navigation labels and page
titles are title case in every locale whose script has case.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- billing.previewChange and billing.change switch between paid personal
  plans in place: an upgrade charges the prorated difference now
  (always_invoice), a downgrade credits unused time (create_prorations),
  both with the previewed proration date. A declined card is a declared
  PAYMENT_REQUIRED; a scheduled cancellation, the same plan, or no
  subscription are declared conflicts. The updated subscription is stored
  the way the webhook stores it.
- billing.cancel and billing.resume: personal, or a team the viewer
  administers. An optional cancel reason goes only to analytics.
- billing.current: the viewer's plan for the account menu and upgrade
  prompts.
- Checkout accepts a validated same-origin /dashboard returnTo; the
  completed checkout returns there with ?welcome=<plan>.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…upgrade prompts

- Plan & billing and each team's Billing tab share one plan picker: the
  current plan marked (with the subscription's own price), Free upgrades
  through Checkout, Pro and Max switch in place behind a preview dialog
  with the prorated amount, the Free card cancels with the end date, what
  stops then, and an optional reason, and a cancelling plan resumes. The
  portal is only for the payment method and invoices.
- Plan & billing is personal by default; ?team= keeps working for old
  links. The Billing scopes list is gone (teams have their own tab).
- After checkout, Plan & billing shows a welcome with first steps; a page
  that sent the user to checkout shows a one-line welcome.
- Cloud, iOS TestFlight, and Mobile devices show a Requires Pro panel to
  Free viewers; Upgrade returns to the same page.
- The account menu shows the plan, and Upgrade for Free accounts.
- All strings in 20 locales.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…granted plan

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: cover SwiftPM warning in package lane tolerance

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: ignore warning text in package lane error check

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ocks (#6443)

* Improve setup prerequisite errors

* Tighten GhosttyKit lock handling

* Clarify Xcode setup preflight

* Sanitize setup error output

* Print sanitized lock mkdir errors

* Register GhosttyKit lock before cleanup

* Harden GhosttyKit lock diagnostics

* Move the Xcode preflight ahead of setup mutations

Main now preflights the Metal toolchain before any setup mutation, and it
already stops a Command Line Tools-only machine. Run one xcodebuild -version
check first so a missing Xcode or an unaccepted license gets a direct message,
and drop the developer-directory path glob, which could reject valid Xcode
installs.

Co-authored-by: archit-goyal <go4archit@gmail.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Retry when a stale GhosttyKit lock disappears

---------

Co-authored-by: Leo Li <cheerleaderleo@outlook.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The subnav read teams from the Hexclave SDK cache, which the dashboard's
own team mutations never refresh, so a deleted team stayed in the Settings
subnav beside every team page. It now reads the team catalog query that
create, rename, leave, and delete already invalidate.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: cover current base ref for registry guard

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: resolve registry guard base ref at runtime

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The current card said "$50/mo" while the others said "$50 per month".
The picker now takes the subscription's Stripe price as data and renders
every card as amount + "per month" (per seat for Team), adding "billed
annually" for yearly prices.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
teamleaderleo and others added 27 commits September 30, 2026 19:02
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* ci: report a cancelled run instead of a routing failure

The `tests` gate in ci.yml and the `ios-tests` gate in test-ios.yml both
run with `if: ${{ always() }}` so that they always report, because both
are in the branch protection required list. A cancelled run therefore
reaches them with every job it depends on reporting `cancelled`, and
their routing check fails with only `changes: cancelled` or
`detect-ios-changes: cancelled` to show for it.

That message describes a routing defect, and the pull request does not
have one. It cost three pull requests a diagnosis each this week before
the answer turned out to be the same every time: rerun the cancelled
run. Reruns of two of them came back green with no change to the branch.

So say that instead. Each gate now has a step that runs only when the
run was cancelled and prints what happened and what clears it.

Both gates still fail on a cancelled run. A cancelled run tested
nothing, and reporting success would let a pull request merge on a green
required check that no suite stood behind. `cancelled()` is true only
for a cancelled run, so a job that timed out still reaches the routing
check and fails there, which is where a timeout belongs.

`cancelled()` is available in `jobs.<job_id>.if` and
`jobs.<job_id>.steps.if` only, not in a step's `env:`, so this is two
mutually exclusive steps rather than one step reading a variable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: detect the cancellation in the gate script, not in a step condition

The first version of this used a second step gated on `cancelled()`. That
only covers a cancelled workflow run: `cancelled()` says nothing about an
individual dependency, so a job cancelled on its own still fell through
to the routing verdict and still reported as a routing defect.

Both gates already read every dependency's result out of `toJSON(needs)`,
so the cancellation is visible there. Checking it in the script covers a
cancelled run and a cancelled dependency with one branch, needs no step
condition, and drops the second step.

The verdict is unchanged in both other cases: a run where everything
reported still exits 0, and a genuine routing failure still exits 1 with
the message it had before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* Add setting actions, setting presets, and cmux config set

cmux.json actions can now change settings: "type": "setting" with a
path and one of set, toggle, cycle, or unset, and "type": "settingPreset"
to apply a named partial settings object from the new top-level
settingPresets. `cmux config get|set|unset|toggle|cycle|preset` drive the
same JSONConfigStore.apply path, so every entrypoint validates against
the schema, keeps comments and unrelated keys, and publishes one atomic
write under the cooperative writer lock.

Setting actions only run when the global cmux.json declares them; project
configs and packs can't ship a button that rewrites global settings.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Address review: empty preset objects, fail-closed trust, docs

- A preset's empty nested object merges nothing instead of replacing
  the whole section; a preset that sets nothing is refused.
- A setting action with no source path fails closed.
- Docs, schema, and comments say packs referenced by the global config
  may declare setting actions (they inherit its source path), and that
  confirm doesn't apply.
- Move the CLI help note below the subcommand list.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Honor confirm on setting actions and explain keys containing "."

A setting or settingPreset action with "confirm": true now asks before
saving and shows the equivalent `cmux config` command. The project-action
trust prompt never covers the global config, so this is the only prompt
these actions get.

A path that only resolves when one of its keys contains "." (for example
a workspaceGroups.byCwd entry for ~/src/app.web) is refused with a
keyContainsDot error that says so, instead of "isn't a cmux setting".
Such keys stay unsupported; there is no escaping syntax. Docs, the CLI
contract, and the cmux-settings skill say so.

Adds a test that packs the global config references may declare setting
actions while a project config's packs may not, and that confirm reaches
both the palette action and the tab bar button.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Write fractional setting values in their short form

JSONSerialization prints a Double with 17 significant digits, so
`cmux config set terminal.scrollSpeed 1.4`, a cycle entry, or a preset
leaf wrote 1.3999999999999999 into cmux.json and the CLI echoed it.
CmuxSettingValue now hands JSONSerialization an NSDecimalNumber built
from Swift's shortest round-trip text, and preset leaves are re-encoded
through CmuxSettingValue so every path writes 1.4.

CI run 36310670056 on e478696 caught this through the new
commandLineDescriptions and confirmationDialogShowsTheEquivalentCommand
tests; this adds a file-text assertion for set and preset.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Start setting changes from the live value; keep project configs out

Review fixes:

- Most settings live in UserDefaults, and cmux.json only overrides them.
  toggle, cycle, and get now read a key the file doesn't set from the
  app's UserDefaults (CmuxSettingLiveValues, backed by the setting
  catalog) before the schema default, so the first press after changing
  a setting in the Settings window flips what the user sees. The app
  reads its own defaults; the CLI reads the enclosing app's domain.
  `cmux config get` says where the value came from, and `unset` is
  described as removing the key from cmux.json.
- A project config could override a global setting action's title,
  shortcut, or confirm through an actions entry or a tab bar button.
  Both overrides are now ignored for setting actions unless they come
  from the global config.
- The failure alert falls back to a localized "couldn't read or save
  cmux.json" message for store errors that have no description.
- keyContainsDot only fires when the rejoined key exists in the file or
  looks like a path, so a typo under a map keyed by names stays an
  unknown path.
- `cmux config get --json` uses `path` and `file` like the other
  subcommands, plus `source`.
- The schema no longer says confirm doesn't apply to setting actions.
- The setting-action docs keys are translated into the 18 other web
  locales, and CHANGELOG has a line.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Only trust live values that are stored in cmux.json form

Second review pass:

- Some UserDefaults values aren't stored the way cmux.json spells them:
  app.minimalMode is a presentation-mode string, and
  app.keepWorkspaceOpenWhenClosingLastSurface is stored as the opposite
  flag. Lists and maps are often stored as text. The live resolver now
  maps those two explicitly and otherwise accepts only a scalar whose
  type the schema allows at that path (CmuxConfigSchemaPathLookup gains
  declaredTypes(at:)). A catalog-wide test fails if a key the resolver
  accepts stores a default that disagrees with the schema default, so a
  new transformed key has to get a mapping.
- Re-setting a fraction that's already in the file is no longer reported
  as a change: receipts and the no-op check compare numbers in the same
  short form they are written in.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Map the session content width sentinel back to false

The catalog test caught terminal.sessionContentMaxWidth: UserDefaults
stores -1 for "no cap", which cmux.json spells false. The live resolver
now maps widths under the minimum to false.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Move the changelog line to the end of Unreleased/Added

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix setting action docs rich text

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix setting presets from global packs

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix settings module import

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix config executor settings import

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add setting actions dogfood tour

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix setting actions tour JSON

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Refine setting actions dogfood tour

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Make setting actions tour reload config

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Isolate setting actions dogfood config

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Use isolated home in setting actions tour shell

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix setting actions tour cleanup

Use the isolated tour home when restoring config so cleanup cannot remove the user's config.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Harden setting action persistence

Prefer current global presets, guard symlink retargets during publication, and keep tour cleanup isolated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: predicted echo stays off for surfaces not known to be remote

A local shell that echoes slower than the threshold would draw predicted
glyphs, and every registered surface's output is copied and scanned while
the setting is on. Both tests fail on main.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Predict echo only on remote terminals; withdraw on pasted and sent input

Every terminal registered with the prediction center, and a local shell
was kept out only by its measured echo latency. A shell on a loaded Mac can
echo slower than the 25 ms threshold, and while the setting was on every
terminal's PTY output was copied and scanned on the main actor. Each
surface is now classified at its first keystroke after prediction starts
for it: only a terminal whose shell runs on another machine (an SSH or
Cloud machine, a remote tmux pane, another signed-in Mac) gets an active
engine, and the output inbox accepts only those surfaces, so a local
terminal's reads are dropped before any copy.

Input that bypassed the keystroke path (a paste completing, text and keys
sent over the socket, the TextBox, a paired phone) left an armed echo run
in place. A key typed right after a paste was drawn where the pasted text
was about to land. That input now withdraws the prediction, like an
editing key does.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Turn predictive local echo on by default and move it to Terminal settings

With prediction limited to terminals whose shell runs on another machine,
the setting leaves Beta Features and defaults on. It is now
Settings > Terminal > Predictive Local Echo and `terminal.predictiveLocalEcho`
in cmux.json and the schema. It stays stored under the former beta
UserDefaults key, so anyone who turned the beta off keeps it off.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Honor the catalog default for predicted echo; withdraw on remote tmux named keys

The center read its setting with UserDefaults.bool(forKey:), which reads an
unset key as off, so a catalog default of on would never reach it. It now
takes the default from the catalog. A transport-named key in a remote tmux
pane skipped the prediction path and left drawn glyphs in place.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Classify Dock terminals for predicted echo; reclassify when a remote session starts or ends

A remote terminal moved into a Dock was classified local because only the
workspace split tree was consulted, so it never predicted. Classification
now resolves the hosting Dock first, as terminal link opening does.

Classification was fixed at a surface's first keystroke, so a terminal typed
into before its remote session registered (a reconnect that keeps the
surface) never predicted in that runtime. A change to a workspace's remote
terminal set now withdraws and classifies the surface again at its next
keystroke.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: deleted text reappears after Backspace, Ctrl-U, Ctrl-W over in-flight predictions

A frame-composing harness (grid plus overlay every 16.7 ms) flags any column
the user deleted and saw blank that later shows text again. 6 of 8 fail.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep deleted predictions blank until the remote erases them

A Backspace over a glyph still in flight no longer lets its echo paint the
character back: the cell is drawn blank until the erase lands. Ctrl-U,
Ctrl-W and Option-Backspace mask in-flight glyphs the same way, and keys
whose effect is unknown (Return, arrows, pastes) raise a barrier behind
the glyphs already sent instead of dropping them, so nothing vanishes
and reappears. A glyph retyped onto a blank cell replaces the blank.

The simulation now checks blanks at the end of the run: a blank over text
already where it belongs fails only if the remote never rewrites that cell.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: Return answered by a newline suspends prediction; Ctrl-U after Return blanks the submitted line

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Resolve the barrier on the remote's newline; match Ctrl-U and Ctrl-W by character

Output answering Return, an arrow or a paste (CR LF, a redraw) resolved the
barrier only when printable; a newline counted as a misprediction, so fast
typing followed by Return could suspend prediction. A barrier that expires
unanswered guessed nothing and no longer counts either. Ctrl-U behind a
barrier acts on the new line and leaves the submitted glyphs drawn.
Ctrl-U and Ctrl-W are recognized by the layout's character rather than the
physical key.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Trace predicted echo in dev builds while a flag file exists

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: drain-before-parse draws echoed glyphs over the cells to their left

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: zsh highlighting suspends, images keep stale glyphs, slow jittery links stop drawing, main-actor scan cost

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: links over 1.5 s never arm; one stall stops drawing until the user pauses

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: type-ahead and masked password prompts draw secrets; status repaints stop prediction

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep predicted echo safe at password prompts, alive through rewrites and stalls

Found while dogfooding over a slow link and by adversarial tests:

- zsh answers the second key of a line with BS and a rewrite of the first
  ("\bec"), which withdrew everything as a misprediction and left prediction
  dark for the rest of the line. The engine now remembers what it saw echoed
  on the row and accepts a move back over it followed by the same text; CSI n D
  reports its count for this. zsh-syntax-highlighting recolouring is the same.
- A key typed ahead of a password prompt got its cooked-mode echo and armed
  the run, so the rest of the password was drawn. Drawing now needs two
  echoes in a row since the run last ended, and a `*` never counts.
- An overdue echo or unmodelled cursor motion now hides what is drawn but
  keeps matching the keystrokes in flight, so one stall no longer leaves the
  line untracked while the user keeps typing. The lifetime scales to three
  round trips on links slower than 500 ms.
- A cursor save and restore (a status line or clock repaint) is ignorable;
  kitty graphics and sixel images are disruptive.
- Output with nothing in flight is skimmed instead of classified byte by
  byte, and the inbox keeps at most 1 MiB per surface, reporting a drop.
- Backspace over an echo not yet painted blanks it until the erase. Blanks
  that output removes stay until the next presented frame, and the host
  drains a surface's pending output before retiring at a frame.

The breaker tests that armed on one echo now arm on two. The two render
race tests are disabled until ghostty tees after the parse.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: point the disabled render-race tests at the ghostty tee fix

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: read the final status outside the expectation

The determinism guard reads a duration literal inside an assertion as a
wall-clock check. This one is the engine's simulated clock.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: a predicted run over the layer-hosting terminal renders nothing

Ghostty makes the terminal view layer-hosting. The overlay drew in
draw(_:), which AppKit called with the host's whole bounds as the dirty
rect, and the run never reached the screen: on a dev build with a 1.8 s
link, ten predicted glyphs sat drawn, unhidden and on top for over a
second with nothing visible.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Render the predicted run into the overlay layer's contents

The terminal view is layer-hosting (ghostty's IOSurface layer is its
layer), and AppKit's draw pass for a subview of it never put the run on
screen. Render the run into a bitmap of the overlay's own size and set it
as the layer's contents from updateLayer, which does not depend on that
pass.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: a layout hold with no presented frame freezes the overlay

On a dev build the rendered-frame callback never fired: it is counted in
the Metal layer's nextDrawable, and ghostty presents through its own
IOSurface layer. The hold set when an erase removed blanks was never
released, so the overlay kept a blank over a cell the grid had already
redrawn, and hid every later prediction on that surface.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Release the layout hold when no frame arrives

The hold waits for a presented frame, and the host may never report one.
Release it after confirmationHold, the same bound a held confirmation
has, and include that deadline in nextExpiry so the host's timer re-anchors
the overlay.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add a predicted-echo recording script for dogfood GIFs

Puts a local sshd behind a delay proxy, opens a cmux ssh workspace on a
tagged build, types paced keys through the debug socket and records the
window as a GIF. The dogfood tour format cannot express a slow link, a
remote workspace, timed keys or video.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix predictive echo for consumed bindings

* pin Ghostty tee fix release

* fix predictive echo binding classification

* classify Ghostty local-only bindings

* fix Ghostty binding flag typing

* record local-only GhosttyKit checksum

* rebase Ghostty binding flag onto main

* pin Ghostty local-only binding fix

* record rebased GhosttyKit checksum

* rebase Ghostty binding fix onto current main

* fix Ghostty gitlink

* record current GhosttyKit checksum

* retain GhosttyKit checksum after main merge

* ci: record rebased GhosttyKit checksum

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(ci): update-homebrew must keep triggering-run metadata out of scripts

update-homebrew.yml substitutes github.event.workflow_run.head_branch and
the dispatch input into its `run:` script as text. The test renders the
version step the way the runner does, with a hostile branch name, and
checks nothing executes; it also requires the gate to accept only the
real release workflow run from a tag push in this repository. Found by a
Codex Security audit (both models).

Red: python3 tests/test_ci_homebrew_untrusted_input.py

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(ci): update-homebrew takes run metadata through env and trusts only the real release run

Makes tests/test_ci_homebrew_untrusted_input.py green (previous commit).

Root cause: the version step substituted
github.event.workflow_run.head_branch and the dispatch input into its
script text, and the gate accepted any run of a workflow with the
release workflow's display name. The values now reach the script as
env vars (INPUT_VERSION, HEAD_BRANCH, EVENT_NAME), and the gate requires
the run to be .github/workflows/release.yml, from this repository, for a
push. Dispatch keeps working as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
#16148)

* ci: survive an owned Mac whose Homebrew prefix the runner does not own

Run 36752535712 built the app-host and UI test product on
cmux-austin-mini-1-glaeda-1, 17 minutes, then lost the whole job in the
"Install tmux" step:

  Error: /opt/homebrew/Cellar is not writable. You should change the
  ownership and permissions of /opt/homebrew/Cellar back to your
  user account:
    sudo chown -R cmux /opt/homebrew/Cellar

brew refuses outright when the prefix belongs to another account, so
the step exits 1, "Resolve selectors against the built tests" is
skipped, and the job reports TEST_RESULT=failed with TEST_SUMMARY and
TEST_OUTPUT both empty. The dispatch looks like a test failure and names
no cause short of reading the raw log.

scripts/ci/brew-ensure.sh retries the install as the prefix owner, and
when even that cannot produce the command it fails with the runner name
and the package to provision. Both brew installs in the E2E action use
it. Each call site keeps its old inline path behind an -x check, because
this action runs against the tested revision's checkout, which may
predate the script.

This does not provision that mini; it stops one missing package from
throwing away a build that already succeeded, and makes the next
occurrence legible from the step summary.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* ci: reach the helper from the workflow revision and classify its failure

Review of the first commit found five ways the fix would not have helped the
run it was written for.

The helper was called workspace-relative, so it came from the tested revision.
test-e2e.yml checks the action out of the workflow's own revision into
.e2e-workflow with a sparse-checkout that did not list the helper, so a
re-dispatch of the ref from the failing run, and every step of a regression
bisect into older history, would have taken the old inline brew install and
failed identically. Add the helper to both sparse-checkout blocks and prefer
the .e2e-workflow copy at both call sites, the way the frames script already
does.

Guard on -f and invoke through bash instead of -x. A lost mode bit made the
ffmpeg branch fall through to an unconditional brew install, which fails on an
unowned prefix even when ffmpeg is already there: worse than no helper at all.

The retry now runs sudo -n, gated on sudo -n true, which is how the rest of
this action and run-in-console-session.sh treat passwordless sudo on these
Macs. Without -n, a host with a controlling tty would block on the prompt
until the job timed out, holding a Mac.

HOMEBREW_NO_AUTO_UPDATE was a command prefix on the first install only, and
sudo's env_reset would drop it regardless, so the retry could trigger a full
brew update inside a step that had already burned the build. Carry it over the
sudo hop with /usr/bin/env, as action.yml:545 does.

Every failure path now prints [cmux-ci machine: brew-provision], so
machine_failure.py classifies it directly instead of depending on Homebrew's
own wording reaching the log. A root-owned prefix gets its own message, since
brew refuses to run as root and no hop fixes that.

Verified: shellcheck and bash -n clean, actionlint clean on test-e2e.yml, both
YAML files parse, tests/test_ci_machine_failure.py passes with the new case (8
tests), tests/test_ci_self_hosted_guard.sh passes, and all eight branches of
the script exercised under brew/stat/sudo shims exit 0 only with the command
present and non-zero only with it missing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(ci): pin the new helper in the e2e sparse-checkout guard

test_the_build_runner_runs_the_tests_so_a_run_queues_once asserts the
sparse-checkout list verbatim, so adding scripts/ci/brew-ensure.sh to it
reddened app-host-execution. The addition is intended: without it the
helper is absent from .e2e-workflow and the action falls back to the
tested revision's copy.

Also assert every entry exists. A sparse-checkout of a path that is not
in the repo is silent, so a typo there would leave the action reading a
missing file and silently taking the older copy instead.

Verified: python3 tests/test_ci_e2e_compilation_cache.py, 25 tests OK.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: cover agent activity after turn completion

* fix: keep agent status running while hooks show activity

* fix: ignore ambiguous late tool activity

* Add subagent and waiting agent states to the sidebar

The compact status glyph had one "running" state, so a pane running a
fan-out of subagents, a pane parked on a background command, and a pane
typing a reply all looked identical. Two of those are worth telling
apart: subagent work is the loudest thing an agent does, and a pane
waiting on a deterministic wakeup is not asking for anything.

Claude's hooks now report what a running pane is running on through a
new `set_status --work=running|subagents|waiting` option:

- PreToolUse with `tool_name` of `Task` reports subagents. A Task call
  blocks the parent inside the tool until its subagents finish, so the
  state holds for exactly that span and the next parent hook clears it.
  No counter to drift.
- Stop with a live background task or scheduled wakeup reports waiting
  instead of running. A re-entrant Stop stays running: that is the agent
  itself still going.

The work state rides alongside the agent lifecycle rather than inside
it. A waiting pane keeps reporting a running lifecycle on purpose, so
hibernation can never SIGTERM live background work; the work state is
presentational only, and the resolver reads it before the lifecycle
branch. Waiting wins only when every agent in the workspace reports it,
so one agent still working keeps the row running.

Glyphs: subagents is a pulsing gray connected-points symbol, waiting is
a still gray hourglass. Waiting does not pulse, because the agent is
parked and a pulsing hourglass would claim otherwise. Both are
configurable through `sidebar.compactStatusIcons`, and both reach the
non-compact rows through the icon the hook sends.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
(cherry picked from commit b4bee23)

* Add dogfood tours for the new sidebar agent work states

Two tours over the same five workspaces (subagents, waiting, running,
needs input, idle): one with the compact glyph on, one with it off so
the metadata rows show the icons the hooks send.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 5904e8a)

* Fix the work-state compile breaks and the shared-key waiting regression

CI on 5904e8a caught two compile breaks the branch shipped with: the two
`shouldReplaceStatusEntry` call sites in SidebarOrderingTests never gained
the new `workState` argument, and a new control-socket test called
`hasPrefix` on an optional response. `everyIconSlotHasADistinctState` also
still pinned 11 icon slots against the 13 the branch now has. The cmuxTests
target could not build, so none of the branch's own tests ran.

The review that ran alongside it found three behavioral defects:

Subagents never appeared on a current Claude Code. The PreToolUse row
matched only `tool_name == "Task"`, and 2.x sends `Agent` for the same
spawn. Both names now count, the way `AgentChatSessionRegistry.isTaskSpawn`
already handles it for the mobile child-run tracker.

An hourglass could cover a pane that was still working. Status entries are
keyed per workspace while lifecycle states are keyed per panel, so two
Claude panes in one workspace share one `claude_code` entry and the second
to report wins. Waiting now also requires that every running lifecycle is
covered by a waiting report, so a sibling pane mid-tool-call keeps the row
running. Two panes both waiting under one key read as running, which is
the conservative direction.

The work state is now listed by `list_status` and `sidebar_state` as
`work=<state>`, so the state behind the glyph is observable instead of
screenshot-only.

Also: the doc comment promised that an unknown work state degrades to a
plain running row, while the socket rejects the whole `set_status` the way
it already rejects an unknown `--format`; the comment now describes what
the code does. `SidebarAgentWorkState.parse` dropped a `_`/`-` pass that
no input could reach and a singular `subagent` alias the socket rejects,
so the two parses accept the same set. The glyph header and
docs/configuration.md listed Running above Waiting while the resolver
checks Waiting first.

Tests: the renamed spawn tool, the two-pane shared-key case both ways, two
agents both parked, the listing line, and a pin on the raw values the
sidebar and control-socket copies of the wire contract share.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 1bad27d)

* Append the work option last so it does not split an older command prefix

CI caught this on the app-host lane:
CLINotifyProcessIntegrationRegressionTests.testClaudePromptSubmitFrom
NewSessionCanReplaceStoppedSession asserts the prompt-submit command as a
prefix through `--tab=`, and `--work=running` was being inserted between
`--color=` and `--tab=`, so the prefix no longer matched. Three assertions
in tests/test_claude_hook_clear_running_status.py use the same contiguous
fragment and would have failed on their own lane for the same reason.

None of those four assertions is about work states; they check that
prompt-submit sets Claude running on the right tab. Options are
order-independent on the wire, since the coordinator reads a parsed option
dictionary, so the new optional one goes at the end of the command instead
and the older assertions stay intact. Updating them to expect
`--work=running` would have coupled four unrelated checks to this feature
and broken them again the next time the work state for prompt-submit
changed.

Pinned by a new test in ClaudeHookWorkStateTests: the running command must
still start with the historical prefix and must end with the work option.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(sidebar): list the work option in the socket help output

The `help` text for `set_status` was the one place that still omitted
`--work`, while the usage and error strings in both coordinator copies
already list it.

Pin the work-state ordering test through the workspace id, so it stands
in byte for byte for the prefix the older suites assert, and say in the
comment why order independence holds: every option here is `--key=value`,
which a future bare flag would not be.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep the sidebar pill on Running for a re-entrant Claude Stop

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: replay deterministic wait tools in sidebar hooks

* fix: show deterministic Claude waits in sidebar

* fix: reopen identityless fresh tool activity

* fix: gate identityless activity by hook phase

* test: model fresh identityless tool resolution

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add the window recording request and frame model to CmuxFoundation

A clip needs its parameters settled before any capture starts: the format
and its default frame rate, the scale and width caps, the crop rectangle in
window points, the file name, and how the frame size follows a window that
is resized mid-clip. None of that needs a window or a screen, so it lives in
CmuxFoundation with tests instead of inside the capture session.

WindowRecordingFrameGeometry aspect-fits each captured image into the frame
size chosen at the start, so a resize letterboxes rather than stretches, and
a crop is adopted once and then kept for the rest of the clip.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add cmux record for capturing a cmux window to mp4 or gif

An agent working in its own cmux can describe what it changed, but it cannot
show it. `cmux record start` films a cmux window, or a region of one, and
writes an mp4 or a gif that can go straight into a pull request. `cmux record
note` drops a caption into the clip as it runs, so a reader can follow what
was being done.

Capture uses ScreenCaptureKit's own-process content, so only cmux's own
windows are reachable and no Screen Recording permission is asked for.
Encoding is AVAssetWriter for mp4 and ImageIO for gif; frame times come from
when each frame was sampled, so playback matches what happened. A recording
stops itself at --max-seconds, and one runs at a time, so an agent that goes
away cannot leave a capture running.

The window.record.* socket methods run on the socket worker and are not
main-thread callable: the window being filmed has to keep drawing while the
sampler runs. They are deliberately absent from the remote relay allowlist,
which stays default-deny, because a clip is local screen content rather than
an object of the peer's workspace.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add failing tests for the window recorder's crash and early-stop paths

A review of the recorder found seven defects the existing tests missed.
These are the reds, before the fixes:

- `--fps nan`, `--fps inf` and `--fps 1e30` reach `Int(Double)` in the
  request decoder and trap, which takes the whole app down with the
  socket. The package suite now crashes with "Double value cannot be
  converted to Int" instead of finishing.
- a gif whose recording stopped long before its frame budget cannot be
  finalized, because the budget is handed to ImageIO as the number of
  images the file will contain, so the documented `--gif` flow leaves no
  file at all.
- `record stop` after a clip reached its own `--max-seconds` limit
  reported "no recording is running" rather than the finished clip.
- an unknown recording id, and a note for a clip that already stopped,
  had no coverage at all.
- the mp4 writer's no-frames path was not checked for leaving a stub
  file behind, and the two-frames-one-tick test only asserted a nonzero
  duration rather than the timestamps it exists to pin down.
- `fit` upscales a window that shrank mid-clip, which its own name says
  it does not do.

`remember` on the registry stops being private so a test can set up a
finished clip without a window on screen.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix the window recorder's crash, race and early-stop paths

Seven defects a review found, in the order they bite a caller:

- `--fps nan` and `--fps 1e30` trapped in `Int(Double)` and took the app
  down. The decoder now refuses a value that is not finite or does not
  fit an `Int`.
- a gif was created with the recording's frame budget as its declared
  image count, so a clip stopped before that budget could not be
  finalized and left no file at all. The count is a floor rather than a
  cap, and how many frames a clip ends with is only known when it stops,
  so declare one and add as many as were captured.
- `stop` could close the file while an append was suspended inside the
  session actor, which is an uncatchable AVFoundation exception for the
  mp4 writer and concurrent state for the gif writer. Appends and the
  close now take turns.
- two `record start` calls could both pass the one-at-a-time check,
  because the check was followed by the await that opens the capture.
  The slot is claimed before that await.
- `record stop` after a clip reached its own `--max-seconds` limit said
  no recording was running; it reports the finished clip, as `status`
  already did.
- the mp4 writer cancelled a writer it had never started when a clip
  captured no frames, which raises rather than returns an error.
- the output file was deleted before the first frame was written, so a
  failed recording destroyed whatever was at `--out`. Frames go to a
  hidden sibling file and only move into place once the clip closes, a
  directory or device at that path is refused, and a partial file is
  cleaned up on every failure path.

Also: a window closed mid-clip reported an empty rectangle and was
captured as a one-pixel frame for the rest of the clip, so an empty
rectangle ends the recording and keeps what came before; `fit` no longer
magnifies a window that shrank mid-clip, which is what its name always
claimed; and `window.record.*` is listed in `system.capabilities` for
local callers, where the relay policy still filters it out.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Refuse a record flag that swallowed the next flag as its value

`cmux record start --label --gif` named the clip "--gif" and recorded an
mp4, because the shared option parser takes whatever follows a flag. The
record command now hands a value that looks like a flag back to the
unexpected-arguments check, which names both, and the positional
recording id for `stop`/`status` no longer accepts one either.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Map an unusable recording output path to invalid_params

`outputNotAFile` was added to `WindowRecordingSessionError` in the fix
commit for the output-path handling, but the router's switch over that
enum was never extended, so the app target stopped compiling:

    Sources/TerminalController+WindowRecording.swift:175:13:
      error: switch must be exhaustive

Nothing on the pull request caught it. The checks that run on a pull
request build the packages and the tooling, not the app target, so this
first appeared in a dispatched UI test run:

    https://github.com/manaflow-ai/cmux/actions/runs/36412249068

The missing case is `invalid_params`: `--out` naming a directory, a
device or anything else the recorder may not replace is the caller's
parameter, not a cmux failure, and an agent that reads `internal_error`
retries instead of fixing its flag.

The mapping had no test at all, which is why a missing case could sit
here. `recordingErrorCode(for:)` is now internal and every case of the
three error types it understands is pinned, including the unrecognized
fallback. A compile error cannot be committed as a red test first, since
the test would not build either; the build log above is the red
evidence, and the test is what keeps the codes from drifting.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Share the recorder's parameter decoding with a screenshot request

`window.screenshot` takes the same region, the same scale and width caps and
the same rule about an output path as `window.record.start`, and a caller who
finds a rectangle in a clip should be able to shoot it with the same four
numbers. Written twice, the two would drift: one would round a region's edges
differently, or accept a relative path, or trap on `--max-width 1e30` where the
other refuses it.

So the decoding moves into `WindowCaptureValueDecoding`, which both requests
translate into their own public failures. The recorder's API and every message
it produces are unchanged; `WindowRecordingRequest.make` now decodes the format
first and maps the shared failure back into its own, which is why the format
still names itself in a path error.

`WindowRecordingFrameGeometry.plan` gains an overload taking a region, a scale
and a width quantum rather than a recording request, so a still can plan its
crop with the code a clip uses. H.264's even-width rounding becomes that
quantum instead of being read off the format.

`WindowRecordingLabel` takes a fallback, so a label that sanitizes away to
nothing is filed as a screenshot rather than under the word "recording", and
`WindowRecordingOutputNaming` gains the screenshot directory the DEBUG command
has always used plus a filename overload taking an extension.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Discard a timed-out record start and keep standalone -- in notes

A record start that timed out on the socket kept running and could
register a recording the caller had been told failed. The registry now
tracks each start by token; on timeout the socket handler abandons it
and waits for the session to release its writer and partial file before
answering. A start that is stopped or abandoned mid-capture no longer
opens a writer, and fail() releases a writer even after the session
reached a terminal state.

record note now strips only a leading -- terminator, and cli.usage.record
carries every catalog locale.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Scope recording limits and output naming onto WindowRecordingRequest

The package conventions lint rejects all-static public types, so the
limits and the default output naming become extensions on the request.
The mp4 test names its encoded clip length plainly so the determinism
check does not read it as a measured wall-clock time.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add `cmux shot` for screenshotting a cmux window

An agent that films a window with `cmux record` still has no way to grab
one frame of it. A clip is the wrong artifact for "what does this sheet
look like now": it is slower to produce, heavier to attach, and a
reviewer cannot read a label off it.

`window.screenshot` captures one of cmux's own windows to a png or a
jpeg, through the same ScreenCaptureKit path the recorder uses. No
Screen Recording permission is involved, so it works on a release build
and inside CI, where the DEBUG-only `screenshot` command does not exist
at all.

The two commands share every decision they could disagree about:
parameter decoding, crop and scale planning, frame capture, output-file
handling and window resolution. `cmux record --region 0,0,420,900` and
`cmux shot --region 0,0,420,900` frame the same rectangle of the same
window, so a detail spotted in a clip can be shot with the same four
numbers.

The image is encoded beside the output path and moved into place, so an
existing file there is replaced only once there is a complete image to
replace it with, and `--out` pointed at a directory comes back as
`invalid_params` rather than deleting it.

Local socket only: `window.screenshot` is not on the `cmux ssh` relay
allowlist, since it reads pixels of windows the remote session does not
own.

## Changelog

Added: `cmux shot` captures a cmux window, or a region of one, to a png
or a jpeg.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Compare recorded pixels through an Equatable value

Optional tuples are not Equatable, so the caption test did not compile.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add a cmux-capture skill for screenshots and clips

`cmux shot` and `cmux record` exist, but an agent looking for a way to show
a reviewer what a change looks like has to find them first. The skill says
which of the two fits a still and which fits a change over time, how to reuse
one `--region` rectangle across both, and what an agent recording its own
cmux needs to know: the recording stops itself, one runs at a time, captions
are drawn in as it goes, and nothing but cmux is in the frames.

`references/commands.md` carries what `--help` cannot: the response fields,
the error codes and the limits the app enforces.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add a capture topic to cmux docs

`cmux docs` is where an agent with no repository checkout looks for a topic,
so the capture skill and its command reference are reachable there too, and
from the agents topic next to the hook integrations.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Guard the capture skill against drift from the code

A doc that names a response field or a state the app no longer emits sends
an agent after something that is not there. Two such errors were in the draft:
it called the recording states `recording` or `stopped` when they are
`recording`, `finished` and `failed`, and called the screenshot `format` field
`png` or `jpg` when the response carries `jpeg` and only the file is named
`.jpg`. Both came from reading the CLI and not the payload.

The guard reads both sides: documented flags against the CLI help, documented
states and formats against their enums, documented response fields against the
keys the payload actually sets, and documented error codes against the ones
the socket methods return. It needs no build, and it runs in the skill
contract workflow when the skill or the code behind it changes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Probe the capture docs topic from the CLI contract

`test_cli_contract_help.py` runs every probe line against the built CLI, so
the new topic is checked the way the dock topic is.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(skills): complete capture discovery contract

* fix(skills): advertise capture docs in CLI help

* Fix window capture edge cases

Validate and safely clamp region geometry, including origin-anchored crops and finite values beyond integer range. Build GIFs only after their exact frame count is known, staging compressed frames on disk. Keep screenshot output promotion on the waiting socket worker so a timed-out capture cannot later overwrite the requested path.

* fix(capture): preserve directories at promotion

* fix(capture): bound frame work and caption history

* fix(capture): isolate uncooperative frame requests

* fix(capture): bound GIF staging and sampling

* Pin the capture doc claims a reviewer found wrong

Nine doc findings from a review of this skill, and three guards so the same
class of drift is caught next time rather than read past.

The corrections: the error table said `invalid_params` covered an unknown record
subcommand, which never reaches the socket; a region under 8 or over 100000
points, a relative `--out` and an `--out` whose extension contradicts the format
were all undocumented; `not_found` covers an unknown recording and a window that
was never open, not only one that closed mid-capture; `conflict` also answers a
`stop` that raced its start; `fps_effective` needs two frames spanning a nonzero
time, not just two frames.

The guards: the documented region bounds are compared against the `Double`
constants both requests enforce, every relative link in the two pages has to
resolve to a file and to a heading that exists, and every Swift file these
checks read has to be in the workflow's trigger paths. The last one was missing
`WindowRecordingRequest.swift`, so a change to the recorder's own bounds could
not have failed this test.

The link guard immediately caught a forward reference: the skill pointed at
`dogfood-scenarios.md#record-a-clip`, a heading that arrives with the dogfood
record step, so the anchor is dropped until that lands.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Wait on deadlines and make the sample schedule a value

Two CI guards refused the capture branch, and both were pointing at
something worth changing rather than at the gate being strict.

`WindowScreenshotWriteTests` waited by yielding a fixed number of times.
N yields is however long N reschedules take, so the wait tightens exactly
when the runner is busy, which is when a passing test turns into a flake.
It now yields until a five second deadline instead.

`WindowRecordingSampleSchedule` was an enum of static functions that took
the previous target back in as a parameter, so the caller held the state
and the type held none. It is now a struct the sampler owns and steps
forward, which is also why its tests can assert on `targetUptime` after a
served slot rather than only on a returned number.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Pin the writer's error codes, which nothing tests any more

`record start` opens the output file before it answers, so an `--out` the
recorder cannot create fails inside the writer rather than in the request
decoder. `TerminalController.recordingErrorCode(for:)` has no branch for
`WindowRecordingWriterError`, so every one of them falls through to
`internal_error`, and an agent told by the docs to fix its flags on
`invalid_params` reads its own unwritable path as a cmux bug and retries.

This test covered it until the branch was repaired; the case that made it
redundant, `stoppedWhileStarting`, is gone, but the writer errors are not.
It is restored first, failing, and the mapping follows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Tell a caller when the recorder could not open their clip

`WindowRecordingWriterError.setup` is the file the caller named refusing to
be created, so it answers `invalid_params`; a frame or a close that fails is
the encoder, so those stay `internal_error`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Pin the limits and codes an agent reads, before fixing them

Three claims the capture docs make that nothing checked.

`invalid_params` tells an agent to edit its flags, so it cannot also mean
"this window has nothing in it": no flag makes an unrendered window
capturable, and the mapper sent all three geometry failures down that
branch. The test now pins the split.

The sample schedule counted one slot past every elapsed slot, so a
sampler arriving exactly on a later slot threw that frame away and waited
a whole interval more for the next one. The new case asks for the slot it
is already in time for.

The limits table was checked for flag names and never for numbers, which
is how it came to promise a gif `--max-width` of 4096 against a ceiling of
1280, and to omit the two gif budgets that have no flag of their own. The
guard now reads the constants and asserts the page states each one, and
the region check compares values instead of mirroring a source line.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep an on-time frame, and stop blaming flags for an empty window

`recordingErrorCode` now answers `not_found` for `emptyWindow`, beside
the two geometry failures a parameter can fix. An agent that reads
`invalid_params` changes its flags and tries again, which no flag would
have helped here.

The schedule rounds up to the first slot at or after now instead of
adding one to the elapsed slots, so a target that is exactly `now` is
served rather than skipped, and it steps once more in the one case where
the division rounds down in binary.

Two comments claimed a window resized mid-clip never ends a recording.
That holds for the frame size, which is fixed when the clip opens, but
not for a `--region` the window has shrunk out of: the crop is planned
again per frame, planning throws, and the clip ends in state `failed`
with the frames it had. Both comments now say so.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Document the gif ceilings, and offer an index in --help

The limits table promised a gif `--max-width` of 4096 while a gif stops
at 1280, and said nothing about the two budgets a gif has and an mp4 does
not: 960 frames, and 4000000 pixels a frame. The reference page owns the
limits and `--help` owns the flag list, so the ceilings go here rather
than into a wider flag column.

Also corrected on that page: a region a shrinking window leaves behind
ends a clip, which the region section had not admitted; `conflict` no
longer covers a stop that raced its start, a case that no longer exists;
an `--out` under a read-only volume is `invalid_params`; a window with no
capturable content is `not_found`; and `rpc window.screenshot` takes a
window id, because the CLI is what resolves a ref or an index. The skill
page said `record status` reports the effective frame rate, which only
`--json` carries.

`--help` for both commands listed `--window <id|ref>` while the usage
line and the reference page both offer an index, and the CLI has always
accepted one. Same token in all nine locales, since it is a spec rather
than prose.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(recording): stop asserting how many samples the encoder writes

`twoFramesInTheSameTickStillGetIncreasingTimes` expected a two frame clip
to read back as exactly two samples, and the CI VMs write six: with no
hardware scaler available, the encoder decides how many samples a clip of
two frames a 1/600 second apart ends up with.

The property under test survives. A dropped frame, which is what an equal
presentation time used to cause, leaves a single sample, so a floor of two
still catches it, and the times still have to start at zero, increase, and
put the second sample a tick rather than a frame interval after the first.
The expectations now print the times they read, so the next surprise from
the encoder arrives with its evidence.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(recording): pin that a gif keeps the frames staged before its limit

A gif that reaches maximumStagedBytes mid-clip currently loses every frame:
finish() retries the frame it could not stage, throws again, and never creates
the destination, so the session deletes the working file. That contradicts
fail(_:), which promises "Keep whatever was captured before the failure".

Red before the fix: the finalized gif does not exist.

## Changelog

none

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(recording): pin what a mid-clip crop replan may and may not fail on

A clip's encoded frame size is fixed by its first frame, so replanning a later
frame with the gif pixel ceiling still applied ends a recording over a size
nothing encodes: a window that grows mid-clip fails with gifFrameTooLarge even
though the frame is drawn into the opening size. A region the window shrank out
of is the one resize that still has to end the clip.

Red before the fix: planCrop does not exist.

## Changelog

none

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(recording): keep a gif's staged frames when the last one cannot stage

`finish()` staged the frame it was holding without clearing `pending`, so a
clip that had already reached its staging limit threw again from the one call
that is supposed to finalize it. The session's `fail(_:)` promises the caller
"whatever was captured before the failure", and that promise was being broken
for exactly the recordings that need it.

Drop the one frame that has no room and finalize the frames before it. A clip
with nothing staged still throws, so a recording with no frames still reports
why it has none.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(recording): re-crop a window that grows mid-clip instead of ending the gif

The per-frame replan called the first-frame planner, which applies the gif
pixel ceiling. A window that grew past that ceiling mid-clip therefore ended
the recording over a frame size nothing encodes: the encoded size was fixed by
the first frame, and the grown crop was only ever going to be drawn into it.

`planCrop` plans the crop with no pixel ceiling, and the session uses it for
later frames. The one resize that still ends a clip is a `region` the window
has shrunk out of, because no rectangle is left to sample.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(cli): give the `screenshot` alias its contract row

`check-cli-contract-verbs` reads dispatched verb strings, and the shot arm
dispatches both `shot` and `screenshot`. Only `shot` had a row, so the guard
failed the whole run with "top-level verb screenshot is in no command table".

The first cell may name a verb and its aliases, so both go in the one row.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: validate screenshot requests and capture docs

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: reproduce answered agent notification ring

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: clear answered agent notification rings

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: clear Codex answered notification rings

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: fence Codex prompt notification retirement

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: normalize agent session fallback matching

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: stamp ordered Codex progress hooks

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: limit Codex progress stamps to safe hooks

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: clear agent rings on terminal input

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: clear agent rings in dock terminals

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: remove unused semantic notification binding

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: isolate agent notification mutation fixtures

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: clear queued mutations between notification fixtures

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: fence superseded agent prompt retirement

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: preserve newer same-session prompts

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: index unread agent prompts

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…15989)

* test: reproduce late Agent Chat startup stealing session focus

* fix: keep late Agent Chat startup replies from changing selection

* Make Agent Chat fork and handoff actions recoverable (#15994)

* test: reproduce stuck Agent Chat fork and handoff lifecycle

* fix: make Agent Chat fork and handoff actions recoverable

* fix(agent-chat): deduplicate reconnecting session actions

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(agent-chat): keep main's pinned tsc in the check script

The main merge resolved agent-chat/package.json toward this branch's
`bun x tsc --noEmit`, but main had since pinned typescript 5.9.2 as a
devDependency and added tests/test_ci_guard_workflow_structure.py, which
asserts the exact check script. CI fast guards failed on that assertion.

The pinned dependency came through the merge, so ./node_modules/.bin/tsc
exists; only the script string was stale.

Verified: python3 tests/test_ci_guard_workflow_structure.py passes, and
`bun run check` in agent-chat is 106 pass / 0 fail across 18 files.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* Checkpoint terminal scrollback so a crash no longer loses it

The 8 s session autosave never captures scrollback, so after a crash or
SIGKILL every terminal restored empty; only clean quit, power-off and
update relaunch persisted it (#2016, #2194).

Add slow, bounded scrollback checkpoints next to the primary snapshot:
- at most every 60 s, only after 5 s without typing, driven from the
  existing autosave timer;
- only terminals that produced PTY output since their last capture,
  flagged by one relaxed atomic load per PTY read in the existing tee;
- at most 3 Ghostty VT exports per checkpoint, one per main-queue turn,
  stopping early on typing or after 50 ms of main-thread capture time;
- truncation, encoding and writes on a utility queue, one file per
  terminal under session-<bundle>-scrollback/, pruned to live panels;
- the same eligibility gates as the quit path (running command,
  hibernated agent), with stale checkpoints deleted.

After an unclean exit, startup restore fills each terminal's missing
scrollback from its checkpoint; the newer of snapshot and checkpoint
wins, and the restore path still applies its own replay gates.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Bound checkpoint capture cost and keep restored scrollback across crashes

Review follow-up for the scrollback checkpoints:

- Split the capture: only Ghostty's VT export runs on main; reading the
  export file, CRLF normalization, the 4000-line tail, truncation,
  encoding and the write run on the utility queue. A terminal whose
  export alone exceeded the 50 ms budget, or failed, is skipped for
  10 minutes.
- Seed checkpoints from restored scrollback when a restore completes,
  keyed by the restored panel ids and without a VT export, so a second
  crash before the next checkpoint keeps it.
- Record scrollbackCapturedAt on snapshots saved with scrollback. A
  newer scrollback-bearing save wins over a checkpoint even when it
  deliberately omitted a terminal's scrollback; the 8 s autosave (no
  marker) still yields to checkpoints.
- Re-mark a terminal pending when its export read or file write fails.
- Keep a terminal pending for one more checkpoint when output arrived
  between planning and capture, since the PTY tee runs before Ghostty
  parses those bytes.
- Disable checkpoints under automated test runs, like session restore.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Discard checkpoint exports on quit and seed only after a crash

- A checkpoint interrupted by quit or restore now deletes the export
  files it already wrote synchronously on main (one unlink each) and
  leaves those terminals pending, instead of handing them to the
  utility queue, which may not run before exit.
- Seed checkpoints from restored scrollback only after an unclean
  previous launch, not on clean launches or manual reopen, to avoid
  rewriting up to 400 KB per terminal needlessly.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix overlapping access in the scrollback checkpoint merge

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Discard checkpoints after a clean exit so a later crash cannot restore them

A clean quit saves scrollback into the snapshot but left the checkpoint
files behind until the next launch's first checkpoint pruned or rewrote
them. Restore reuses snapshot panel ids (workspace and dock terminals),
so if that next launch crashed first, its 8 s autosave (no scrollback,
no capture marker) matched the old records and the crash restore
replayed the earlier launch's scrollback over what the quit saved.

Startup now deletes the checkpoint directory when the previous launch
exited cleanly, before the restore decision so a skipped restore also
drops them. After an unclean exit the checkpoints are that launch's own
and are still merged and kept until the restored terminals are seeded.

Moves the checkpoint coordinator setup into the AppDelegate checkpoint
extension to keep AppDelegate.swift's growth down.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Inject one scrollback checkpoint activity owner and delete removals before returning

The PTY tee bridge, the checkpoint coordinator and the persist step now share
one `TerminalScrollbackCheckpointActivity` owned by the composition root
(`GhosttyApp.terminalScrollbackCheckpointActivity`) instead of a `.shared`
singleton, so a caller cannot hand the coordinator a different instance than
the one the tee writes.

Planned removals (a panel that stopped being eligible) are now applied with
`queue.sync` after earlier queued writes, and only captures and pruning go
async. A crash before the async block ran used to leave the old checkpoint on
disk for the next unclean restore to merge.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test scrollback checkpoint output interleaving

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Make scrollback checkpoints generation-safe

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Preserve scrollback for idle agent sessions

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Match checkpoint agent liveness policy

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: cover Ctrl-letter key event encoding

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: encode synthetic Ctrl-letter keys

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: compare named Ctrl-letter bytes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: report named key releases

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: avoid shadowing Ghostty key event type

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: retain synthetic key press capture

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: match Ghostty legacy Ctrl-I and Ctrl-M

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: restore bonsplit pin for macOS compile

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
#15420)

* Move browser and tmux OSC filter tests into their package test targets

BrowserSystemProxyMirror, BrowserUserAgentPolicy, BrowserNavigationDecisionHandler,
BrowserHiddenWebViewDiscardManager and BrowserHTTPBasicAuthPromptCoordinator live in
CmuxBrowser, and RemoteTmuxNotificationOSCFilter in CmuxRemoteSession (#13787), but
their 68 tests compiled into cmuxTests and ran inside the launched app host. They use
only package API, so they move unchanged except for dropping the cmux import.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Drop the blank line the removed import block left

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep app shard metadata aligned with package tests

* test: give split fixtures deterministic geometry

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add failing tests for agent wake verification

The tests need new types, so this commit also adds minimal stubs:
AgentWakeVerificationState never changes state, and the Workspace entry
points (beginAgentWakeVerification, failAgentWakeVerification,
dismissAgentWakeFailure) do nothing. The state machine and Workspace
tests fail against these stubs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Check that a woken agent came back

Waking a hibernated agent types its resume command into a fresh shell,
but nothing checked whether the agent started. A missing launcher or a
gone session left the pane at a shell prompt with no sign of failure.

A wake now starts a check for the pane. An agent hook reporting a PID or
a lifecycle state counts as success. The check fails when the resume
command returns to the prompt first, or when nothing reports within 90
seconds and no live agent process is found. A failure shows a banner on
the pane (Retry, Show command with Copy, Close), an "Agent didn't
resume" sidebar row, and one notification feed entry. The row is kept
out of the saved session. Closing the pane, hibernating it again, or a
later agent report clears the check.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Only fail quick resume exits and never retype into a running program

A resume command that returns after 20 seconds ran the agent, which the
user then quit; for agents without hooks that was reported as a failed
wake. It now counts as resumed. Retry is hidden, and refused, while a
command still runs in the pane, since the text would go to that program.
The failed-wake row survives a sidebar reset, and a retired workspace
skips the deadline.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Confirm wakes only from the woken agent's own evidence

Hook reports now count only when they come from the woken agent and
name their pane explicitly; the focused-pane fallback no longer confirms
a wake. A resume command that ran a while is no longer taken as proof:
a pending check samples the pane for a live agent process every few
seconds instead, and a command that ends before any confirmation fails
the wake. The verification deadline constants are nonisolated, which
fixes the new Swift warning.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(cloud): cover the reads that find a Cloud machine gone

A status read and an access preflight both move the row to `destroyed`
themselves when the provider no longer has the machine. Once either does,
`destroyVm` can never see the row again (its lookup skips destroyed rows)
and the provider-status cron skips it too, so the model-plane revoke and
the `vm.destroyed` usage event that the cron performs for the same
transition have to happen at these write sites as well.

Fails today: both retire the row and record nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): finish the destroy work where a machine is first found gone

Three places retire a Cloud machine's row after the provider says it no
longer has the machine: the status read, an access operation's resume
preflight, and the stats read. Each wrote `status = destroyed` and stopped
there, while the reconcile cron doing the same transition also revoked the
machine's model-plane tokens and recorded a `vm.destroyed` ledger event.

That difference was permanent, not a race to lose. `destroyed` is terminal:
`findUserVm` hides such a row from every destroy request and
`reconciliationCandidates` drops it from the cron, so whichever of the three
got there first left a machine that never appears as destroyed in the ledger
and never has its route tokens marked revoked, with nothing able to finish
the job afterwards.

All three now go through one `applyObservedProviderStatus` that performs the
write and, when the write lands on `destroyed`, the same revoke and ledger
event as the cron. The status route hands `getVm` the model-plane revoker it
already builds for delete. The new `provider_status_*` destroy reasons join
the analytics allowlist so the event says which read noticed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): derive the observed status inside the shared destroy path

Review found that the previous commit took the row's new status from each
caller. That let the access preflight and the stats read hardcode
"destroyed" for a provider 404, including for a machine with a persistent
home volume, which observedDbStatus maps to "paused" because the compute
is gone and the machine is not. Those rows were terminalized and billed a
vm.destroyed that never happened, and nothing revisits a terminal row to
take either back.

The status is now derived inside applyObservedProviderStatus, so every
entrypoint agrees about what a 404 means. reopenBaseIfProviderDeleted is
the one caller that must override it, and passes forceStatus with the
reason: its row is a Base's active generation, and leaving it paused would
hand the same dead provider id back on every later open.

Also threads the model-plane revoker through getVmStats from its route,
covers the stats entrypoint including the home-volume case, and fixes the
two test-file regressions the previous commit shipped.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* vm: say what a paused 404 row means for credentials

The access preflight carried a comment claiming a revoke was pointless
there because the credentials were already inert. That is false for the
case this branch introduces: authenticateRouteToken and
authenticateVmAuthorization both accept `paused`, so route tokens on a
volume-backed machine stay valid after this write, for the rest of their
30-day lifetime.

Keep the behavior, which is what getVm and the reconcile cron already do
for the same observation, and replace the comment with what is true. Also
drop the stale "next fleet refresh drops it" line, correct "seven call
sites" to eight, stop asserting a resurrection path that is not
implemented, and type usageEventSource as VmDestroySource so an unknown
source cannot silently degrade in PostHog.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(cloud): reproduce lost destroy ledger retry

* fix(cloud): retire missing provider compute durably

* fix(cloud): durably drain observed destroy cleanup

* fix(cloud): keep destroy cleanup queue fair and indexed

* fix(vm): make observed cleanup retries index-safe

* test(cloud): reproduce account deletion cleanup loss

* fix(cloud): retain pending destroy cleanup on account deletion

* fix(cloud): transfer account cleanup to durable outbox

* fix(cloud): reject malformed cleanup outbox rows

* fix(cloud): validate cleanup outbox shape

* Persist explicit destroy cleanup for retry

* fix cloud observed cleanup retries and concurrent migration

* fix migration retry for invalid concurrent indexes

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat: surface configured actions and launchers

* fix(actions): preserve surface button action identity

* fix(actions): report exact placements and global edit target

* test(actions): pin exact surface placement identity

* test(config): preserve surface action references

* fix(actions): drop hidden surface action references

* test(config): drop references for hidden surface buttons

* fix: localize configured actions empty state

* fix: sync upstream CmuxFoundation import repair (#13236)

Applies upstream 968cef2 to unblock this PR's app-host compilation.

* fix(actions): localize discovery metadata labels

* fix: localize actions discovery titles and open button

Give the Actions menu item and dialog their own whole-string keys instead of
borrowing the titlebar-layout debug label and concatenating, and label the
open button with the existing Open cmux.json key rather than a raw path.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: document French identity for the actions discovery titles

"Actions" is spelled the same in French, and cmux.json is a file name.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat: optionally isolate native agent scratch files

* fix(settings): register canonical agent scratch JSON path

* fix(settings): document canonical agent scratch in schema

* build(schema): regenerate embedded config schema

* docs(settings): add canonical scratch key reference

* fix: localize canonical scratch schema description

* fix: regenerate config schema after current-main refresh
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
App-host test processes on one runner share the app's standard defaults.
Every main-window close writes its frame there, and the next process opens
its launch window from that frame. A 320-point fixture closed by an earlier
process therefore sized the launch window, every createMainWindow() copied
it, and split admission (#15392) refused the side-by-side splits the
equalize, font-size, remote tmux and manual-unread tests need.

A test process now forgets the persisted window frame at launch, next to the
existing shortcut resets for the same shared profile.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On macOS 15 runners a titled NSButton has an extra NSTextField descendant,
so walking every descendant found 3 text fields for rows with an action
button and failed the labels.count == 2 check. The row's labels are its
direct subviews; AppKit sizes the button's internal text.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…15921)

The blocked-cancellation test held the transport writer only after a fake
daemon event arrived, so the writer had to be held within the 1 s attach
RPC deadline. When that round trip took longer on a loaded runner, the
cancellation write ran first, succeeded, cancelled the write deadline, and
the expected transport stop never came (line 182 timed out after 10 s).

The test now holds the writer first and then runs the handler a timed-out
pty.attach invokes, keeping the writer held until the test ends. The
companion test still covers the timeout-to-cancellation wiring.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Austin Wang <austinwang115@gmail.com>
* Add iOS Feed tab with X-style rows and inline agent actions

A new Feed tab, distinct from Notifications, mirrors the Mac's workstream
Feed on the phone as an X-style full-width timeline: agent output renders
inline (preambles, plans, prompts, tool lines) and every respondable
interaction is actionable in place — permission requests (allow once/always/
all/bypass/deny), exit-plan approvals with inline revise feedback, question
option chips with an Other… free-text lane, and turn-complete replies routed
to the owning terminal.

Mac side: a `feed.v1` mobile capability exposes `feed.list` (workstream
items enriched with resolved workspace/surface targets, display titles, and
conversation context) plus `feed.permission.reply`, `feed.question.reply`,
and `feed.exit_plan.reply`, which dispatch to the same handlers the control
socket uses so every entrypoint resolves through FeedCoordinator.deliverReply.
WorkstreamStore gains a monotonic revision with a change hook that emits
revision-only `feed.changed` events to subscribed phones.

iOS side: per-Mac revision-guarded snapshots aggregated across paired Macs,
optimistic local resolution reconciled by refresh, capability-gated targets,
and lifecycle shared with the notification feed. Turn replies reuse
mobile.terminal.input. Strings ship EN+JA in the CmuxMobileShellUI catalog.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Feed round 2: never render wire JSON, keep the Feed decision-dense

Dogfood feedback: raw JSON must never reach the screen, and routine tool
churn buries the rows a user can act on.

Mac: feed.list now parses the ExitPlanMode envelope to plan text (same
parse as the Mac Feed panel) and drops PreToolUse/PostToolUse telemetry
except failed tool results, so the 200-row budget carries notable rows.

iOS: tool payloads humanize to command/file_path/prompt fields or
key: value pairs and never render braces; plan bodies defensively parse
older Macs' JSON envelopes; the view filters residual tool churn from
older hosts; question options render as X-style stacked poll bars (the
custom wrapping Layout is gone) and a pending question always offers the
Other… free-text lane even when its prompt failed to parse.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Fix shadowed normalized() helper in tool-line humanizer

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Lower the keyboard after inline feed replies

Send drops the field's focus, and swiping the feed dismisses the
keyboard interactively, so an abandoned reply never pins it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Feed round 3: toolbar filter menu, reply markers, stable expand, no title

The filter moves from an in-list segmented control to the same toolbar
menu control the Workspaces list uses (Mail-style icon, filled while
narrowing). Replying to a turn-complete row now records the reply on
that row: a quote-reference line names the message being answered and a
"You" bubble shows what was sent, persisted across snapshot refreshes.
Show more expands without animation so the top of the row stays fixed
while the text grows downward. The Feed drops its navigation title.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Restore main's submodule pins clobbered by merge resolution

The conflict-resolution git add recorded the stale local ghostty and
bonsplit checkouts as gitlinks, reverting main's pin bumps (bonsplit
c5452560 carries TabDragTransferRegistration.clearResidualCapability,
which Sources/SessionDragSessionSource.swift requires).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Feed: X-style reply composer sheet; the timeline never hosts a keyboard

Every free-text lane (turn-complete reply, question Other…, plan
Revise…) now opens one compose sheet modeled on X's reply flow: the
quoted message with a thread line, Cancel top-left, a prominent
Reply/Send capsule top-right, auto-focused editor, and swipe-down or
Cancel to dismiss — so the keyboard always has an obvious way down.
Stop rows show a quiet bubble Reply button in place of the old inline
field.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Feed: no divider above the top row

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Feed design pass: prompt rows are authored by You; tighter collapse

Sim UX pass findings: user-prompt rows read backwards ('Claude received
your prompt') — X-style, your prompt is your post, so they now render a
person avatar, 'You' as the author, and 'prompted <agent>' as the
headline. Long outputs collapse at 360 chars / 8 lines.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Feed: unwrap double-encoded plan envelopes instead of rendering them

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Feed: the user's own prompts are not standalone Feed events

Dogfood: the timeline was a wall of 'You prompted …' rows. A prompt is
not news to its author — it already surfaces as the quoted context line
under the agent rows it produced. Filter userPrompt rows on the Mac
(freeing the row budget for notable rows) and on the phone (for older
hosts).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Feed round 4: one action-button family, swipe triage, visible Reply

Dogfood feedback on the controls. Approve/Allow, Revise/Always, and
Deny now share one option-bar-shaped family — accent fill, quiet fill,
and red-tinted fill — replacing the stock bordered mix with its bare
red-on-gray Deny; the overflow ellipsis renders at full strength in a
matching chip. Rows gain mark-read-style leading swipe triage (Done /
Needs Input) backed by a local override that adjusts the badge, filter,
and unread dot without touching the row's answerable state. The stop
row's Reply is an accent-tinted button instead of a quiet footnote, and
the composer's Reply loses its custom capsule so the toolbar's own pill
is the only background.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Fix mobile feed replies and duplicate stop rows

* Use inline reply action styling

* Test Feed reply routing and real terminal submission

* Cover submit-key selection after switching agent providers

* Share Feed prompt submission and honor the active agent

* Require asynchronous prompt submission on the control socket

* Reproduce Enter arriving inside agent paste protection

* Separate agent paste and submit input before acknowledging replies

* Cover provider identity through Feed ingestion

* Preserve Grok identity in Feed events

* fix: remove stale Sentry test target linkage

* Verify Feed submissions against real provider editors

* Improve provider verification harness

* Document Feed session scope and outstanding verification

* Cover refreshing an existing personal development account

* Allow verified personal credential refresh and record connection blocker

* Record tagged dev sign-in gate result

* Record working developer pairing and main refresh scope

* Update readiness fixture for installed bundle evidence

* Cover awaiting notification reply delivery before acknowledgement

* Await terminal submission in notification delivery callers

* Verify notification delivery with refreshed Feed builds

* Keep Feed timestamp formatting on its row model

* Fix main API rename in legacy mobile host

* Adapt feed target resolution to main

* Record main rebuild and queued phone install

* Record transport main refresh

* Record refreshed builds and remaining phone auth blockers

* Test full Feed completion preservation and paged reading

* Open retained Feed messages in a full-text reader

* Exercise Feed reader retry scrolling and dismissal in iOS UI test

* Run focused Feed reading tests through the registered hosted workflow

* docs: explain Feed data flow and record delivery and new scope

* test: require Feed-scoped search and conditional inline expansion

* fix: inline Feed expansion, scoped search, shared controls and event navigation

* fix: isolate Feed search navigation and apply exact computer scope

* fix: contain inline touch targets and seed navigation tests through store state

* fix: handle Feed search in the workspace preview scaffold

* Document Feed navigation verification and queued phone update

* Test unified Feed open action

* Use one Open action for Feed destinations

* Persist feed replies and mirror notifications

* Keep empty notification feed revisions compatible

* Document durable notification feed inputs

* Preserve durable feed order on replay

* Cover exact feed event reply identity

* Record unified feed open action

* Persist feed generation across restarts

* docs: point feed data flow to current source

* docs: record current feed build status

* Add swipeable multi-question Feed answers

* Render Feed event content as Markdown

* Fix Markdown intent raw value conversion

* Fix Markdown sheet navigation modifiers

* Test Markdown preview truncation and expansion offsets

* Preserve rendered Markdown while truncating Feed previews

* Keep question answer composition on its draft type

* Test block Markdown rendering in Feed previews

* Render block Markdown in Feed row previews

* Record Feed Markdown verification and paired build delivery

* Test Feed first row spacing beneath the toolbar

* Remove the empty large-title area above Feed rows

* Use shallow checkout for manually dispatched iOS checks

* Record Feed spacing proof and authenticated phone delivery

* Keep current Feed delivery status consistent in scope notes

* Quiet the Feed Reply action to a text-scale secondary control

The filled blue arrow and 44-point layout frame made Reply louder than
the row content and added a gap above it. Render an outline icon and
footnote label in secondary color, keep blue for links, and extend only
the hit area to 44 points so row layout stays compact.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Record the quieter Feed Reply action in scope notes

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add iMessage-style Feed quote bubbles behind a CMUX Labs switch

Quoted prompts render as outlined bubbles with a leading tail, and a
recorded reply shows the answered message as an outlined quote above a
filled reply bubble, matching iMessage inline replies. CMUX Labs >
Feed Bubble Quotes switches between this and the original leading-bar
quotes; it defaults on in DEBUG and is unavailable in production.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Use the Messages tail geometry for Feed quote bubbles

The first tail flattened the bottom-leading corner into a slab. Draw the
body inset by the tail width and hook the tail out and back into the
bottom edge, as Messages does, so stroked and filled bubbles read as
iMessage bubbles.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Record Feed bubble quotes Labs switch in scope notes

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Point Feed bubble quotes of the user's words to the right

A quoted prompt is the user's own message, so draw it as a sent bubble:
trailing-aligned, accent-tinted, with the tail mirrored to the right. In
a recorded reply the answered agent text stays a gray leading quote and
the user's reply becomes a filled accent bubble on the trailing side.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Smooth the Feed bubble tail join into the bottom edge

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Record right-pointing Feed quote bubbles in scope notes

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add failing test for tapping links in Feed previews

Feed previews render Markdown links but their text view ignores touches,
so tapping a link such as a PR number does nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Open Markdown links tapped in Feed previews

The preview text view is non-interactive so row gestures keep working,
which also swallowed link taps. A tap recognizer on the container now
begins only when the touch lands on rendered link text, resolves the
destination from the laid-out glyph, and opens it through SwiftUI's
openURL like the Feed's other Markdown text. Other taps pass through.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Show each Feed event's workspace, with an optional tab setting

Finished-turn rows drop the "finished a turn" verb so the author line
reads agent · workspace. Every row shows the workspace it came from
(its title, else the cwd folder). Settings > Display > Show Tab in Feed
appends the event's tab for people who organize agents by tab.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Open Feed rows inside the Feed tab like Notifications rows

Tapping a Feed row now opens the event's workspace or tab, and the
workspace pushes onto the Feed tab's own navigation stack (new
agentFeed deep-link origin), so Back returns to the Feed instead of
jumping to Workspaces. Buttons, links, and See more keep their own
taps; VoiceOver gets an Open action, and rows get stable identifiers.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add Feed offline banner and a Retry action

When the Feed cannot refresh but has cached rows, a banner above them
says it is offline or that a Mac needs an update, matching the
Notifications tab, so stale activity is not mistaken for live. The
full-screen unavailable state gains a Try Again button.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add a CMUX Labs switch to let the Feed replace Notifications

Feed Replaces Notifications hides the Notifications tab on iPhone and
its destination in the iPad sidebar, moving a selected Notifications
tab to the Feed, so the replacement can be dogfooded before the tab is
removed. DEBUG-only; production keeps the Notifications tab.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Pin Xcode and route forks in the Feed reply provider workflow

Both macOS jobs now select Xcode through scripts/select-ci-xcode.sh and
fall back to a GitHub-hosted runner outside manaflow-ai, satisfying the
pinned-Xcode and fork runner routing guards.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Show failed Feed replies and keep the text for retry

A terminal reply that fails now leaves a red note on its row instead of
only a log line. When the owning Mac was offline the note says the
reply was not sent; when the request reached the Mac or its fate is
unknown it says delivery could not be confirmed and offers Open
Terminal, so the user checks before retrying and the text is never
typed twice unseen. Try Again reopens the composer with the text.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Order Workstream log appends and let callers await them

Each append was an independent Task, so two mutations of one item could
reach the JSONL log out of order and a restart would replay the older
version. Appends now form one ordered chain, and flushPersistence()
awaits it. The restart tests use it instead of sleeping 50 ms, which
the test-determinism guard rejects.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Record the overnight Feed readiness pass in scope notes

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Record Feed usage events like the Notifications tab does

Appends agent Feed app events (opened, closed, filter changed, item
opened, reply succeeded, reply failed with its delivery state) to the
diagnostic taxonomy and records them from the Feed store view, so
replacing the Notifications tab does not lose usage visibility. Events
carry only counts, never content.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add a UI test that a Feed row tap opens its destination

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Record Feed analytics and local test results in scope notes

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Record overnight CI findings and the Feed UI test lane gap

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix question pager clipping, nav buttons, and duplicate notification rows

A paged TabView takes its proposed height and centers overflow, so tall
question pages clipped mid-glyph at 212 points. The pager now measures
each page and sizes the frame to the current page. Previous and Next
join the row's shared action-button family (quiet fill and accent fill,
trailing chevron) and navigation is never disabled; only Submit gates
on complete answers. The Feed's actionable banner records its history
notification under the feed event's id, so the mobile Feed's existing
id dedupe drops the duplicate Notification row beside the real one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Keep agent-hook notification echoes out of the mobile Feed

Agent hooks deliver a notification banner beside every feed event, and
the mobile Feed's history merge re-listed those banners as their own
rows, so the phone showed a Notification row duplicating the question,
permission, plan, or turn-completion row beside it. Notifications now
record whether an agent context produced them (threaded through history
records and session restore, optional so old data still decodes), and
the mobile Feed merge skips agent-produced records; human-facing
notifications such as cmux notify and watchers keep their rows.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Record F46 pager fix and F47 echo suppression in scope notes

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Extend the session notification snapshot init with isAgentEvent

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Order snapshot init arguments to match the declaration

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Record overnight verification results and new gaps in scope notes

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Exclude unflagged notification history from the mobile Feed merge

Records written before the isAgentEvent flag existed are mostly agent
echoes, so requiring an explicit human-facing flag removes the stored
Notification rows the user still saw; every notification remains in the
Notifications tab.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Show only the user's reply bubble, matched to quote size, with a soft tail

The reply marker repeated the agent's message as a gray bubble right
under the same text in the row; bubbles now belong to user text alone,
so the marker is just the reply, at the quoted-prompt bubble's footnote
size and padding. The bubble tail ends in a lifted rounded tip instead
of a sharp point. Scope records G22: codex emits one turn through two
lanes with different reason texts, so the equal-reason stop dedupe
keeps both rows.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Unread-by-default Feed rows and one row per provider turn

Every new event now arrives needing the user (dot, badge, and Needs
Input filter) until a visible, explicit interaction clears it: an
answer, a reply, opening the row, reading full text, or the swipe
triage. Scrolling past marks nothing. Read state and its baseline
persist phone-locally. Stops that describe the same turn through two
provider lanes (a truncated single-line preview beside the full text)
collapse onto one row keeping the fuller text within a two-minute
window, fixing the duplicate codex response rows.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Trim the read-row set with a prefix rebuild

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Hide the Notifications tab by default; Feed replaces it

The replacement ships for everyone: the Notifications tab and its iPad
sidebar destination stay hidden unless Settings > Display > Legacy
Notifications Tab brings them back, keeping the legacy screen hidden
but accessible. The former DEBUG-only Labs switch becomes this shipping
setting with its default flipped.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Expand the replied-to event to its full text in the composer

The reply sheet focuses one event, so its quote now renders Markdown
and offers See more when truncated: the full text loads through the
same paged reader the full-text view uses and the sheet scrolls it in
place, with a retry label on failure.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Truncate Feed previews with one layout pass instead of a binary search

Scrolling re-measures rows constantly, and each truncated row paid a
binary search of full TextKit layouts to find its cut point. A single
layout pass now reads where the visible lines end, then steps back
through composed characters until the See more suffix fits, removing
the repeated relayouts behind the reported scroll hitches.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Fix review findings in the unread, dedupe, and notification-flag paths

The feedReplacesNotifications key moves out of the DEBUG block so
Release builds compile. Read state for stops is also keyed by turn, so
the same-turn collapse cannot resurrect a read row when the fuller
lane replaces it, and read keys evict oldest-first through an ordered
list instead of arbitrary Set order. Same-turn matching now requires
the shorter reason to carry the truncation ellipsis, so two distinct
complete reasons never merge. A failed feed fetch leaves its Mac
pending so an equal re-emitted revision still refetches. The truncated
preview gets a maximumNumberOfLines backstop, and notification
snapshots with unknown provenance restore as agent-produced so legacy
banners cannot re-enter the mobile Feed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Fix mobile feed full text layout

* Fix mobile feed reply expansion and empty rows

* Fix feed UI test scope

* Fix async feed notification tests

* Polish mobile feed decision controls

* Fix mobile feed rendering and scroll performance

* Add mobile feed interaction coverage

* test(ios): keep feed UI test build actor-safe

* fix(ios): use supported feed control sizing

* test(ios): expose paged feed questions as containers

* test(ios): swipe the feed pager directly

* test(ios): cover full text inside feed replies

* test(ios): make full reply fixture addressable

* fix(ios): order full reply fixture arguments

* test(ios): verify feed pager intrinsic layout

* fix(ios): measure paged feed questions after sizing

* fix(ios): top align paged feed questions

* Make feed link hit testing layout-safe

* Fix CI cancellation guard assertion

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: austinywang <austinwang115@gmail.com>
Co-authored-by: Lawrence Chen <54008264+lawrencecchen@users.noreply.github.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Oct 6, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: Repository: manaflow-ai/cmux/.coderabbit.yaml
  • Review profile: ASSERTIVE
  • Plan: Advanced
  • Run ID: 62d53544-9dfd-4a1a-9dfa-13a81b00c25b

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants