Repository navigation
Make agent key hints clickable in terminal panes - #17765
Closed
azooz2003-bit wants to merge 166 commits into
Closed
azooz2003-bit wants to merge 166 commits into
azooz2003-bit wants to merge 166 commits into
Conversation
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Behind the new agentActions.keyHints setting (off by default), a key hint that Claude Code, Codex, or OpenCode prints in a terminal pane, such as "ctrl+o to expand" or "esc to interrupt", presses those keys when clicked. Hovering a hint underlines it, shows a pointing hand, and a tooltip names the keys. When the agent has captured the mouse, the click belongs to the agent and only a Command-click presses the hint; a Command-click on a hint then wins over the link and path fallback for that cell. What gets sent respects both binding layers. For Claude Code, a hint printed with its default chord is remapped to the user's chord when ~/.claude/keybindings.json binds that action elsewhere (read on click, cached by modification date). If the pane's Ghostty keybindings consume the chord, the key's legacy terminal encoding is sent instead so the binding does not run. Part of #15223. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…/agent-key-hints # Conflicts: # cmux.xcodeproj/project.pbxproj
…s_binding AgentKeyHintDetector, AgentKeyHintChordResolver and AgentKeyHintClickPolicy become values with their inputs as stored properties, and the legacy key encoding becomes a String initializer, so the package namespace-type lint passes. Package tests link the Ghostty runtime stubs, which now define ghostty_surface_key_is_binding. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…turn Ctrl+Shift+C, ctrl+alt+c and ctrl+shift+z still send 0x03 and 0x1a but normalize to chords outside the blocked set, and "return" reads prose such as "then return to continue the loop" as an Enter hint. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Read exactly the pointer's viewport row (TerminalTextRegion.viewportRow) instead of bottom-aligning the trimmed, unwrapped viewport text, which mapped a click on a blank row to a line of content above it. - Hover reads that one row per new cell, reads the setting once per event, and reuses a keybindings file check for up to 5 seconds. - Press a clicked hint once the double-click interval passes without a second press, so a double or triple click presses nothing. - Bare keys (Enter, Esc, Tab, arrows, letters, Home, End, page keys) count only in the agent's live region: viewport at the bottom, row no more than 8 rows above the cursor. Ctrl and Alt chords count anywhere. "return" is no longer a key word. - Block any Ctrl chord on c, d, z or backslash whatever Shift or Alt is added, in the detector and the keybindings resolver. - Clear the hover on scroll and layout, and re-resolve it when the viewport under a shown hint changes. - Prefer a running or waiting agent over a stale idle lifecycle. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat: make Codex goal resume clickable Add a narrow, live-row allowlist for the exact /goal resume prompt. A click sends the command and Enter through the existing deferred agent hint path. * feat: show hover affordance for Codex action commands
…gation Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…e navigation A pathless hub route puts Plan & billing and every team page beside the Settings subnav (Account, Billing, Teams) while /dashboard/billing and /dashboard/teams/* keep their public URLs. The sidebar keeps one Account entry, Settings, current on every hub page. Navigation labels and page titles are title case in every locale whose script has case. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- billing.previewChange and billing.change switch between paid personal plans in place: an upgrade charges the prorated difference now (always_invoice), a downgrade credits unused time (create_prorations), both with the previewed proration date. A declined card is a declared PAYMENT_REQUIRED; a scheduled cancellation, the same plan, or no subscription are declared conflicts. The updated subscription is stored the way the webhook stores it. - billing.cancel and billing.resume: personal, or a team the viewer administers. An optional cancel reason goes only to analytics. - billing.current: the viewer's plan for the account menu and upgrade prompts. - Checkout accepts a validated same-origin /dashboard returnTo; the completed checkout returns there with ?welcome=<plan>. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…upgrade prompts - Plan & billing and each team's Billing tab share one plan picker: the current plan marked (with the subscription's own price), Free upgrades through Checkout, Pro and Max switch in place behind a preview dialog with the prorated amount, the Free card cancels with the end date, what stops then, and an optional reason, and a cancelling plan resumes. The portal is only for the payment method and invoices. - Plan & billing is personal by default; ?team= keeps working for old links. The Billing scopes list is gone (teams have their own tab). - After checkout, Plan & billing shows a welcome with first steps; a page that sent the user to checkout shows a one-line welcome. - Cloud, iOS TestFlight, and Mobile devices show a Requires Pro panel to Free viewers; Upgrade returns to the same page. - The account menu shows the plan, and Upgrade for Free accounts. - All strings in 20 locales. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…granted plan Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: cover SwiftPM warning in package lane tolerance Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: ignore warning text in package lane error check Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ocks (#6443) * Improve setup prerequisite errors * Tighten GhosttyKit lock handling * Clarify Xcode setup preflight * Sanitize setup error output * Print sanitized lock mkdir errors * Register GhosttyKit lock before cleanup * Harden GhosttyKit lock diagnostics * Move the Xcode preflight ahead of setup mutations Main now preflights the Metal toolchain before any setup mutation, and it already stops a Command Line Tools-only machine. Run one xcodebuild -version check first so a missing Xcode or an unaccepted license gets a direct message, and drop the developer-directory path glob, which could reject valid Xcode installs. Co-authored-by: archit-goyal <go4archit@gmail.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Retry when a stale GhosttyKit lock disappears --------- Co-authored-by: Leo Li <cheerleaderleo@outlook.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The subnav read teams from the Hexclave SDK cache, which the dashboard's own team mutations never refresh, so a deleted team stayed in the Settings subnav beside every team page. It now reads the team catalog query that create, rename, leave, and delete already invalidate. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: cover current base ref for registry guard Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: resolve registry guard base ref at runtime Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The current card said "$50/mo" while the others said "$50 per month". The picker now takes the subscription's Stripe price as data and renders every card as amount + "per month" (per seat for Team), adding "billed annually" for yearly prices. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* ci: report a cancelled run instead of a routing failure
The `tests` gate in ci.yml and the `ios-tests` gate in test-ios.yml both
run with `if: ${{ always() }}` so that they always report, because both
are in the branch protection required list. A cancelled run therefore
reaches them with every job it depends on reporting `cancelled`, and
their routing check fails with only `changes: cancelled` or
`detect-ios-changes: cancelled` to show for it.
That message describes a routing defect, and the pull request does not
have one. It cost three pull requests a diagnosis each this week before
the answer turned out to be the same every time: rerun the cancelled
run. Reruns of two of them came back green with no change to the branch.
So say that instead. Each gate now has a step that runs only when the
run was cancelled and prints what happened and what clears it.
Both gates still fail on a cancelled run. A cancelled run tested
nothing, and reporting success would let a pull request merge on a green
required check that no suite stood behind. `cancelled()` is true only
for a cancelled run, so a job that timed out still reaches the routing
check and fails there, which is where a timeout belongs.
`cancelled()` is available in `jobs.<job_id>.if` and
`jobs.<job_id>.steps.if` only, not in a step's `env:`, so this is two
mutually exclusive steps rather than one step reading a variable.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* ci: detect the cancellation in the gate script, not in a step condition
The first version of this used a second step gated on `cancelled()`. That
only covers a cancelled workflow run: `cancelled()` says nothing about an
individual dependency, so a job cancelled on its own still fell through
to the routing verdict and still reported as a routing defect.
Both gates already read every dependency's result out of `toJSON(needs)`,
so the cancellation is visible there. Checking it in the script covers a
cancelled run and a cancelled dependency with one branch, needs no step
condition, and drops the second step.
The verdict is unchanged in both other cases: a run where everything
reported still exits 0, and a genuine routing failure still exits 1 with
the message it had before.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* Add setting actions, setting presets, and cmux config set cmux.json actions can now change settings: "type": "setting" with a path and one of set, toggle, cycle, or unset, and "type": "settingPreset" to apply a named partial settings object from the new top-level settingPresets. `cmux config get|set|unset|toggle|cycle|preset` drive the same JSONConfigStore.apply path, so every entrypoint validates against the schema, keeps comments and unrelated keys, and publishes one atomic write under the cooperative writer lock. Setting actions only run when the global cmux.json declares them; project configs and packs can't ship a button that rewrites global settings. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Address review: empty preset objects, fail-closed trust, docs - A preset's empty nested object merges nothing instead of replacing the whole section; a preset that sets nothing is refused. - A setting action with no source path fails closed. - Docs, schema, and comments say packs referenced by the global config may declare setting actions (they inherit its source path), and that confirm doesn't apply. - Move the CLI help note below the subcommand list. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Honor confirm on setting actions and explain keys containing "." A setting or settingPreset action with "confirm": true now asks before saving and shows the equivalent `cmux config` command. The project-action trust prompt never covers the global config, so this is the only prompt these actions get. A path that only resolves when one of its keys contains "." (for example a workspaceGroups.byCwd entry for ~/src/app.web) is refused with a keyContainsDot error that says so, instead of "isn't a cmux setting". Such keys stay unsupported; there is no escaping syntax. Docs, the CLI contract, and the cmux-settings skill say so. Adds a test that packs the global config references may declare setting actions while a project config's packs may not, and that confirm reaches both the palette action and the tab bar button. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Write fractional setting values in their short form JSONSerialization prints a Double with 17 significant digits, so `cmux config set terminal.scrollSpeed 1.4`, a cycle entry, or a preset leaf wrote 1.3999999999999999 into cmux.json and the CLI echoed it. CmuxSettingValue now hands JSONSerialization an NSDecimalNumber built from Swift's shortest round-trip text, and preset leaves are re-encoded through CmuxSettingValue so every path writes 1.4. CI run 36310670056 on e478696 caught this through the new commandLineDescriptions and confirmationDialogShowsTheEquivalentCommand tests; this adds a file-text assertion for set and preset. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Start setting changes from the live value; keep project configs out Review fixes: - Most settings live in UserDefaults, and cmux.json only overrides them. toggle, cycle, and get now read a key the file doesn't set from the app's UserDefaults (CmuxSettingLiveValues, backed by the setting catalog) before the schema default, so the first press after changing a setting in the Settings window flips what the user sees. The app reads its own defaults; the CLI reads the enclosing app's domain. `cmux config get` says where the value came from, and `unset` is described as removing the key from cmux.json. - A project config could override a global setting action's title, shortcut, or confirm through an actions entry or a tab bar button. Both overrides are now ignored for setting actions unless they come from the global config. - The failure alert falls back to a localized "couldn't read or save cmux.json" message for store errors that have no description. - keyContainsDot only fires when the rejoined key exists in the file or looks like a path, so a typo under a map keyed by names stays an unknown path. - `cmux config get --json` uses `path` and `file` like the other subcommands, plus `source`. - The schema no longer says confirm doesn't apply to setting actions. - The setting-action docs keys are translated into the 18 other web locales, and CHANGELOG has a line. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Only trust live values that are stored in cmux.json form Second review pass: - Some UserDefaults values aren't stored the way cmux.json spells them: app.minimalMode is a presentation-mode string, and app.keepWorkspaceOpenWhenClosingLastSurface is stored as the opposite flag. Lists and maps are often stored as text. The live resolver now maps those two explicitly and otherwise accepts only a scalar whose type the schema allows at that path (CmuxConfigSchemaPathLookup gains declaredTypes(at:)). A catalog-wide test fails if a key the resolver accepts stores a default that disagrees with the schema default, so a new transformed key has to get a mapping. - Re-setting a fraction that's already in the file is no longer reported as a change: receipts and the no-op check compare numbers in the same short form they are written in. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Map the session content width sentinel back to false The catalog test caught terminal.sessionContentMaxWidth: UserDefaults stores -1 for "no cap", which cmux.json spells false. The live resolver now maps widths under the minimum to false. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Move the changelog line to the end of Unreleased/Added Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fix setting action docs rich text Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fix setting presets from global packs Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fix settings module import Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fix config executor settings import Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Add setting actions dogfood tour Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fix setting actions tour JSON Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Refine setting actions dogfood tour Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Make setting actions tour reload config Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Isolate setting actions dogfood config Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Use isolated home in setting actions tour shell Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fix setting actions tour cleanup Use the isolated tour home when restoring config so cleanup cannot remove the user's config. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Harden setting action persistence Prefer current global presets, guard symlink retargets during publication, and keep tour cleanup isolated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: predicted echo stays off for surfaces not known to be remote
A local shell that echoes slower than the threshold would draw predicted
glyphs, and every registered surface's output is copied and scanned while
the setting is on. Both tests fail on main.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Predict echo only on remote terminals; withdraw on pasted and sent input
Every terminal registered with the prediction center, and a local shell
was kept out only by its measured echo latency. A shell on a loaded Mac can
echo slower than the 25 ms threshold, and while the setting was on every
terminal's PTY output was copied and scanned on the main actor. Each
surface is now classified at its first keystroke after prediction starts
for it: only a terminal whose shell runs on another machine (an SSH or
Cloud machine, a remote tmux pane, another signed-in Mac) gets an active
engine, and the output inbox accepts only those surfaces, so a local
terminal's reads are dropped before any copy.
Input that bypassed the keystroke path (a paste completing, text and keys
sent over the socket, the TextBox, a paired phone) left an armed echo run
in place. A key typed right after a paste was drawn where the pasted text
was about to land. That input now withdraws the prediction, like an
editing key does.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Turn predictive local echo on by default and move it to Terminal settings
With prediction limited to terminals whose shell runs on another machine,
the setting leaves Beta Features and defaults on. It is now
Settings > Terminal > Predictive Local Echo and `terminal.predictiveLocalEcho`
in cmux.json and the schema. It stays stored under the former beta
UserDefaults key, so anyone who turned the beta off keeps it off.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Honor the catalog default for predicted echo; withdraw on remote tmux named keys
The center read its setting with UserDefaults.bool(forKey:), which reads an
unset key as off, so a catalog default of on would never reach it. It now
takes the default from the catalog. A transport-named key in a remote tmux
pane skipped the prediction path and left drawn glyphs in place.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Classify Dock terminals for predicted echo; reclassify when a remote session starts or ends
A remote terminal moved into a Dock was classified local because only the
workspace split tree was consulted, so it never predicted. Classification
now resolves the hosting Dock first, as terminal link opening does.
Classification was fixed at a surface's first keystroke, so a terminal typed
into before its remote session registered (a reconnect that keeps the
surface) never predicted in that runtime. A change to a workspace's remote
terminal set now withdraws and classifies the surface again at its next
keystroke.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: deleted text reappears after Backspace, Ctrl-U, Ctrl-W over in-flight predictions
A frame-composing harness (grid plus overlay every 16.7 ms) flags any column
the user deleted and saw blank that later shows text again. 6 of 8 fail.
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Keep deleted predictions blank until the remote erases them
A Backspace over a glyph still in flight no longer lets its echo paint the
character back: the cell is drawn blank until the erase lands. Ctrl-U,
Ctrl-W and Option-Backspace mask in-flight glyphs the same way, and keys
whose effect is unknown (Return, arrows, pastes) raise a barrier behind
the glyphs already sent instead of dropping them, so nothing vanishes
and reappears. A glyph retyped onto a blank cell replaces the blank.
The simulation now checks blanks at the end of the run: a blank over text
already where it belongs fails only if the remote never rewrites that cell.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: Return answered by a newline suspends prediction; Ctrl-U after Return blanks the submitted line
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Resolve the barrier on the remote's newline; match Ctrl-U and Ctrl-W by character
Output answering Return, an arrow or a paste (CR LF, a redraw) resolved the
barrier only when printable; a newline counted as a misprediction, so fast
typing followed by Return could suspend prediction. A barrier that expires
unanswered guessed nothing and no longer counts either. Ctrl-U behind a
barrier acts on the new line and leaves the submitted glyphs drawn.
Ctrl-U and Ctrl-W are recognized by the layout's character rather than the
physical key.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Trace predicted echo in dev builds while a flag file exists
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: drain-before-parse draws echoed glyphs over the cells to their left
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: zsh highlighting suspends, images keep stale glyphs, slow jittery links stop drawing, main-actor scan cost
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: links over 1.5 s never arm; one stall stops drawing until the user pauses
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: type-ahead and masked password prompts draw secrets; status repaints stop prediction
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Keep predicted echo safe at password prompts, alive through rewrites and stalls
Found while dogfooding over a slow link and by adversarial tests:
- zsh answers the second key of a line with BS and a rewrite of the first
("\bec"), which withdrew everything as a misprediction and left prediction
dark for the rest of the line. The engine now remembers what it saw echoed
on the row and accepts a move back over it followed by the same text; CSI n D
reports its count for this. zsh-syntax-highlighting recolouring is the same.
- A key typed ahead of a password prompt got its cooked-mode echo and armed
the run, so the rest of the password was drawn. Drawing now needs two
echoes in a row since the run last ended, and a `*` never counts.
- An overdue echo or unmodelled cursor motion now hides what is drawn but
keeps matching the keystrokes in flight, so one stall no longer leaves the
line untracked while the user keeps typing. The lifetime scales to three
round trips on links slower than 500 ms.
- A cursor save and restore (a status line or clock repaint) is ignorable;
kitty graphics and sixel images are disruptive.
- Output with nothing in flight is skimmed instead of classified byte by
byte, and the inbox keeps at most 1 MiB per surface, reporting a drop.
- Backspace over an echo not yet painted blanks it until the erase. Blanks
that output removes stay until the next presented frame, and the host
drains a surface's pending output before retiring at a frame.
The breaker tests that armed on one echo now arm on two. The two render
race tests are disabled until ghostty tees after the parse.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: point the disabled render-race tests at the ghostty tee fix
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: read the final status outside the expectation
The determinism guard reads a duration literal inside an assertion as a
wall-clock check. This one is the engine's simulated clock.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: a predicted run over the layer-hosting terminal renders nothing
Ghostty makes the terminal view layer-hosting. The overlay drew in
draw(_:), which AppKit called with the host's whole bounds as the dirty
rect, and the run never reached the screen: on a dev build with a 1.8 s
link, ten predicted glyphs sat drawn, unhidden and on top for over a
second with nothing visible.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Render the predicted run into the overlay layer's contents
The terminal view is layer-hosting (ghostty's IOSurface layer is its
layer), and AppKit's draw pass for a subview of it never put the run on
screen. Render the run into a bitmap of the overlay's own size and set it
as the layer's contents from updateLayer, which does not depend on that
pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: a layout hold with no presented frame freezes the overlay
On a dev build the rendered-frame callback never fired: it is counted in
the Metal layer's nextDrawable, and ghostty presents through its own
IOSurface layer. The hold set when an erase removed blanks was never
released, so the overlay kept a blank over a cell the grid had already
redrawn, and hid every later prediction on that surface.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Release the layout hold when no frame arrives
The hold waits for a presented frame, and the host may never report one.
Release it after confirmationHold, the same bound a held confirmation
has, and include that deadline in nextExpiry so the host's timer re-anchors
the overlay.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add a predicted-echo recording script for dogfood GIFs
Puts a local sshd behind a delay proxy, opens a cmux ssh workspace on a
tagged build, types paced keys through the debug socket and records the
window as a GIF. The dogfood tour format cannot express a slow link, a
remote workspace, timed keys or video.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix predictive echo for consumed bindings
* pin Ghostty tee fix release
* fix predictive echo binding classification
* classify Ghostty local-only bindings
* fix Ghostty binding flag typing
* record local-only GhosttyKit checksum
* rebase Ghostty binding flag onto main
* pin Ghostty local-only binding fix
* record rebased GhosttyKit checksum
* rebase Ghostty binding fix onto current main
* fix Ghostty gitlink
* record current GhosttyKit checksum
* retain GhosttyKit checksum after main merge
* ci: record rebased GhosttyKit checksum
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(ci): update-homebrew must keep triggering-run metadata out of scripts update-homebrew.yml substitutes github.event.workflow_run.head_branch and the dispatch input into its `run:` script as text. The test renders the version step the way the runner does, with a hostile branch name, and checks nothing executes; it also requires the gate to accept only the real release workflow run from a tag push in this repository. Found by a Codex Security audit (both models). Red: python3 tests/test_ci_homebrew_untrusted_input.py Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(ci): update-homebrew takes run metadata through env and trusts only the real release run Makes tests/test_ci_homebrew_untrusted_input.py green (previous commit). Root cause: the version step substituted github.event.workflow_run.head_branch and the dispatch input into its script text, and the gate accepted any run of a workflow with the release workflow's display name. The values now reach the script as env vars (INPUT_VERSION, HEAD_BRANCH, EVENT_NAME), and the gate requires the run to be .github/workflows/release.yml, from this repository, for a push. Dispatch keeps working as before. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
#16148) * ci: survive an owned Mac whose Homebrew prefix the runner does not own Run 36752535712 built the app-host and UI test product on cmux-austin-mini-1-glaeda-1, 17 minutes, then lost the whole job in the "Install tmux" step: Error: /opt/homebrew/Cellar is not writable. You should change the ownership and permissions of /opt/homebrew/Cellar back to your user account: sudo chown -R cmux /opt/homebrew/Cellar brew refuses outright when the prefix belongs to another account, so the step exits 1, "Resolve selectors against the built tests" is skipped, and the job reports TEST_RESULT=failed with TEST_SUMMARY and TEST_OUTPUT both empty. The dispatch looks like a test failure and names no cause short of reading the raw log. scripts/ci/brew-ensure.sh retries the install as the prefix owner, and when even that cannot produce the command it fails with the runner name and the package to provision. Both brew installs in the E2E action use it. Each call site keeps its old inline path behind an -x check, because this action runs against the tested revision's checkout, which may predate the script. This does not provision that mini; it stops one missing package from throwing away a build that already succeeded, and makes the next occurrence legible from the step summary. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * ci: reach the helper from the workflow revision and classify its failure Review of the first commit found five ways the fix would not have helped the run it was written for. The helper was called workspace-relative, so it came from the tested revision. test-e2e.yml checks the action out of the workflow's own revision into .e2e-workflow with a sparse-checkout that did not list the helper, so a re-dispatch of the ref from the failing run, and every step of a regression bisect into older history, would have taken the old inline brew install and failed identically. Add the helper to both sparse-checkout blocks and prefer the .e2e-workflow copy at both call sites, the way the frames script already does. Guard on -f and invoke through bash instead of -x. A lost mode bit made the ffmpeg branch fall through to an unconditional brew install, which fails on an unowned prefix even when ffmpeg is already there: worse than no helper at all. The retry now runs sudo -n, gated on sudo -n true, which is how the rest of this action and run-in-console-session.sh treat passwordless sudo on these Macs. Without -n, a host with a controlling tty would block on the prompt until the job timed out, holding a Mac. HOMEBREW_NO_AUTO_UPDATE was a command prefix on the first install only, and sudo's env_reset would drop it regardless, so the retry could trigger a full brew update inside a step that had already burned the build. Carry it over the sudo hop with /usr/bin/env, as action.yml:545 does. Every failure path now prints [cmux-ci machine: brew-provision], so machine_failure.py classifies it directly instead of depending on Homebrew's own wording reaching the log. A root-owned prefix gets its own message, since brew refuses to run as root and no hop fixes that. Verified: shellcheck and bash -n clean, actionlint clean on test-e2e.yml, both YAML files parse, tests/test_ci_machine_failure.py passes with the new case (8 tests), tests/test_ci_self_hosted_guard.sh passes, and all eight branches of the script exercised under brew/stat/sudo shims exit 0 only with the command present and non-zero only with it missing. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(ci): pin the new helper in the e2e sparse-checkout guard test_the_build_runner_runs_the_tests_so_a_run_queues_once asserts the sparse-checkout list verbatim, so adding scripts/ci/brew-ensure.sh to it reddened app-host-execution. The addition is intended: without it the helper is absent from .e2e-workflow and the action falls back to the tested revision's copy. Also assert every entry exists. A sparse-checkout of a path that is not in the repo is silent, so a typo there would leave the action reading a missing file and silently taking the older copy instead. Verified: python3 tests/test_ci_e2e_compilation_cache.py, 25 tests OK. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: cover agent activity after turn completion * fix: keep agent status running while hooks show activity * fix: ignore ambiguous late tool activity * Add subagent and waiting agent states to the sidebar The compact status glyph had one "running" state, so a pane running a fan-out of subagents, a pane parked on a background command, and a pane typing a reply all looked identical. Two of those are worth telling apart: subagent work is the loudest thing an agent does, and a pane waiting on a deterministic wakeup is not asking for anything. Claude's hooks now report what a running pane is running on through a new `set_status --work=running|subagents|waiting` option: - PreToolUse with `tool_name` of `Task` reports subagents. A Task call blocks the parent inside the tool until its subagents finish, so the state holds for exactly that span and the next parent hook clears it. No counter to drift. - Stop with a live background task or scheduled wakeup reports waiting instead of running. A re-entrant Stop stays running: that is the agent itself still going. The work state rides alongside the agent lifecycle rather than inside it. A waiting pane keeps reporting a running lifecycle on purpose, so hibernation can never SIGTERM live background work; the work state is presentational only, and the resolver reads it before the lifecycle branch. Waiting wins only when every agent in the workspace reports it, so one agent still working keeps the row running. Glyphs: subagents is a pulsing gray connected-points symbol, waiting is a still gray hourglass. Waiting does not pulse, because the agent is parked and a pulsing hourglass would claim otherwise. Both are configurable through `sidebar.compactStatusIcons`, and both reach the non-compact rows through the icon the hook sends. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> (cherry picked from commit b4bee23) * Add dogfood tours for the new sidebar agent work states Two tours over the same five workspaces (subagents, waiting, running, needs input, idle): one with the compact glyph on, one with it off so the metadata rows show the icons the hooks send. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> (cherry picked from commit 5904e8a) * Fix the work-state compile breaks and the shared-key waiting regression CI on 5904e8a caught two compile breaks the branch shipped with: the two `shouldReplaceStatusEntry` call sites in SidebarOrderingTests never gained the new `workState` argument, and a new control-socket test called `hasPrefix` on an optional response. `everyIconSlotHasADistinctState` also still pinned 11 icon slots against the 13 the branch now has. The cmuxTests target could not build, so none of the branch's own tests ran. The review that ran alongside it found three behavioral defects: Subagents never appeared on a current Claude Code. The PreToolUse row matched only `tool_name == "Task"`, and 2.x sends `Agent` for the same spawn. Both names now count, the way `AgentChatSessionRegistry.isTaskSpawn` already handles it for the mobile child-run tracker. An hourglass could cover a pane that was still working. Status entries are keyed per workspace while lifecycle states are keyed per panel, so two Claude panes in one workspace share one `claude_code` entry and the second to report wins. Waiting now also requires that every running lifecycle is covered by a waiting report, so a sibling pane mid-tool-call keeps the row running. Two panes both waiting under one key read as running, which is the conservative direction. The work state is now listed by `list_status` and `sidebar_state` as `work=<state>`, so the state behind the glyph is observable instead of screenshot-only. Also: the doc comment promised that an unknown work state degrades to a plain running row, while the socket rejects the whole `set_status` the way it already rejects an unknown `--format`; the comment now describes what the code does. `SidebarAgentWorkState.parse` dropped a `_`/`-` pass that no input could reach and a singular `subagent` alias the socket rejects, so the two parses accept the same set. The glyph header and docs/configuration.md listed Running above Waiting while the resolver checks Waiting first. Tests: the renamed spawn tool, the two-pane shared-key case both ways, two agents both parked, the listing line, and a pin on the raw values the sidebar and control-socket copies of the wire contract share. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> (cherry picked from commit 1bad27d) * Append the work option last so it does not split an older command prefix CI caught this on the app-host lane: CLINotifyProcessIntegrationRegressionTests.testClaudePromptSubmitFrom NewSessionCanReplaceStoppedSession asserts the prompt-submit command as a prefix through `--tab=`, and `--work=running` was being inserted between `--color=` and `--tab=`, so the prefix no longer matched. Three assertions in tests/test_claude_hook_clear_running_status.py use the same contiguous fragment and would have failed on their own lane for the same reason. None of those four assertions is about work states; they check that prompt-submit sets Claude running on the right tab. Options are order-independent on the wire, since the coordinator reads a parsed option dictionary, so the new optional one goes at the end of the command instead and the older assertions stay intact. Updating them to expect `--work=running` would have coupled four unrelated checks to this feature and broken them again the next time the work state for prompt-submit changed. Pinned by a new test in ClaudeHookWorkStateTests: the running command must still start with the historical prefix and must end with the work option. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sidebar): list the work option in the socket help output The `help` text for `set_status` was the one place that still omitted `--work`, while the usage and error strings in both coordinator copies already list it. Pin the work-state ordering test through the workspace id, so it stands in byte for byte for the prefix the older suites assert, and say in the comment why order independence holds: every option here is `--key=value`, which a future bare flag would not be. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Keep the sidebar pill on Running for a re-entrant Claude Stop Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: replay deterministic wait tools in sidebar hooks * fix: show deterministic Claude waits in sidebar * fix: reopen identityless fresh tool activity * fix: gate identityless activity by hook phase * test: model fresh identityless tool resolution --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add the window recording request and frame model to CmuxFoundation
A clip needs its parameters settled before any capture starts: the format
and its default frame rate, the scale and width caps, the crop rectangle in
window points, the file name, and how the frame size follows a window that
is resized mid-clip. None of that needs a window or a screen, so it lives in
CmuxFoundation with tests instead of inside the capture session.
WindowRecordingFrameGeometry aspect-fits each captured image into the frame
size chosen at the start, so a resize letterboxes rather than stretches, and
a crop is adopted once and then kept for the rest of the clip.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add cmux record for capturing a cmux window to mp4 or gif
An agent working in its own cmux can describe what it changed, but it cannot
show it. `cmux record start` films a cmux window, or a region of one, and
writes an mp4 or a gif that can go straight into a pull request. `cmux record
note` drops a caption into the clip as it runs, so a reader can follow what
was being done.
Capture uses ScreenCaptureKit's own-process content, so only cmux's own
windows are reachable and no Screen Recording permission is asked for.
Encoding is AVAssetWriter for mp4 and ImageIO for gif; frame times come from
when each frame was sampled, so playback matches what happened. A recording
stops itself at --max-seconds, and one runs at a time, so an agent that goes
away cannot leave a capture running.
The window.record.* socket methods run on the socket worker and are not
main-thread callable: the window being filmed has to keep drawing while the
sampler runs. They are deliberately absent from the remote relay allowlist,
which stays default-deny, because a clip is local screen content rather than
an object of the peer's workspace.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add failing tests for the window recorder's crash and early-stop paths
A review of the recorder found seven defects the existing tests missed.
These are the reds, before the fixes:
- `--fps nan`, `--fps inf` and `--fps 1e30` reach `Int(Double)` in the
request decoder and trap, which takes the whole app down with the
socket. The package suite now crashes with "Double value cannot be
converted to Int" instead of finishing.
- a gif whose recording stopped long before its frame budget cannot be
finalized, because the budget is handed to ImageIO as the number of
images the file will contain, so the documented `--gif` flow leaves no
file at all.
- `record stop` after a clip reached its own `--max-seconds` limit
reported "no recording is running" rather than the finished clip.
- an unknown recording id, and a note for a clip that already stopped,
had no coverage at all.
- the mp4 writer's no-frames path was not checked for leaving a stub
file behind, and the two-frames-one-tick test only asserted a nonzero
duration rather than the timestamps it exists to pin down.
- `fit` upscales a window that shrank mid-clip, which its own name says
it does not do.
`remember` on the registry stops being private so a test can set up a
finished clip without a window on screen.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Fix the window recorder's crash, race and early-stop paths
Seven defects a review found, in the order they bite a caller:
- `--fps nan` and `--fps 1e30` trapped in `Int(Double)` and took the app
down. The decoder now refuses a value that is not finite or does not
fit an `Int`.
- a gif was created with the recording's frame budget as its declared
image count, so a clip stopped before that budget could not be
finalized and left no file at all. The count is a floor rather than a
cap, and how many frames a clip ends with is only known when it stops,
so declare one and add as many as were captured.
- `stop` could close the file while an append was suspended inside the
session actor, which is an uncatchable AVFoundation exception for the
mp4 writer and concurrent state for the gif writer. Appends and the
close now take turns.
- two `record start` calls could both pass the one-at-a-time check,
because the check was followed by the await that opens the capture.
The slot is claimed before that await.
- `record stop` after a clip reached its own `--max-seconds` limit said
no recording was running; it reports the finished clip, as `status`
already did.
- the mp4 writer cancelled a writer it had never started when a clip
captured no frames, which raises rather than returns an error.
- the output file was deleted before the first frame was written, so a
failed recording destroyed whatever was at `--out`. Frames go to a
hidden sibling file and only move into place once the clip closes, a
directory or device at that path is refused, and a partial file is
cleaned up on every failure path.
Also: a window closed mid-clip reported an empty rectangle and was
captured as a one-pixel frame for the rest of the clip, so an empty
rectangle ends the recording and keeps what came before; `fit` no longer
magnifies a window that shrank mid-clip, which is what its name always
claimed; and `window.record.*` is listed in `system.capabilities` for
local callers, where the relay policy still filters it out.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Refuse a record flag that swallowed the next flag as its value
`cmux record start --label --gif` named the clip "--gif" and recorded an
mp4, because the shared option parser takes whatever follows a flag. The
record command now hands a value that looks like a flag back to the
unexpected-arguments check, which names both, and the positional
recording id for `stop`/`status` no longer accepts one either.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Map an unusable recording output path to invalid_params
`outputNotAFile` was added to `WindowRecordingSessionError` in the fix
commit for the output-path handling, but the router's switch over that
enum was never extended, so the app target stopped compiling:
Sources/TerminalController+WindowRecording.swift:175:13:
error: switch must be exhaustive
Nothing on the pull request caught it. The checks that run on a pull
request build the packages and the tooling, not the app target, so this
first appeared in a dispatched UI test run:
https://github.com/manaflow-ai/cmux/actions/runs/36412249068
The missing case is `invalid_params`: `--out` naming a directory, a
device or anything else the recorder may not replace is the caller's
parameter, not a cmux failure, and an agent that reads `internal_error`
retries instead of fixing its flag.
The mapping had no test at all, which is why a missing case could sit
here. `recordingErrorCode(for:)` is now internal and every case of the
three error types it understands is pinned, including the unrecognized
fallback. A compile error cannot be committed as a red test first, since
the test would not build either; the build log above is the red
evidence, and the test is what keeps the codes from drifting.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Share the recorder's parameter decoding with a screenshot request
`window.screenshot` takes the same region, the same scale and width caps and
the same rule about an output path as `window.record.start`, and a caller who
finds a rectangle in a clip should be able to shoot it with the same four
numbers. Written twice, the two would drift: one would round a region's edges
differently, or accept a relative path, or trap on `--max-width 1e30` where the
other refuses it.
So the decoding moves into `WindowCaptureValueDecoding`, which both requests
translate into their own public failures. The recorder's API and every message
it produces are unchanged; `WindowRecordingRequest.make` now decodes the format
first and maps the shared failure back into its own, which is why the format
still names itself in a path error.
`WindowRecordingFrameGeometry.plan` gains an overload taking a region, a scale
and a width quantum rather than a recording request, so a still can plan its
crop with the code a clip uses. H.264's even-width rounding becomes that
quantum instead of being read off the format.
`WindowRecordingLabel` takes a fallback, so a label that sanitizes away to
nothing is filed as a screenshot rather than under the word "recording", and
`WindowRecordingOutputNaming` gains the screenshot directory the DEBUG command
has always used plus a filename overload taking an extension.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Discard a timed-out record start and keep standalone -- in notes
A record start that timed out on the socket kept running and could
register a recording the caller had been told failed. The registry now
tracks each start by token; on timeout the socket handler abandons it
and waits for the session to release its writer and partial file before
answering. A start that is stopped or abandoned mid-capture no longer
opens a writer, and fail() releases a writer even after the session
reached a terminal state.
record note now strips only a leading -- terminator, and cli.usage.record
carries every catalog locale.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Scope recording limits and output naming onto WindowRecordingRequest
The package conventions lint rejects all-static public types, so the
limits and the default output naming become extensions on the request.
The mp4 test names its encoded clip length plainly so the determinism
check does not read it as a measured wall-clock time.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add `cmux shot` for screenshotting a cmux window
An agent that films a window with `cmux record` still has no way to grab
one frame of it. A clip is the wrong artifact for "what does this sheet
look like now": it is slower to produce, heavier to attach, and a
reviewer cannot read a label off it.
`window.screenshot` captures one of cmux's own windows to a png or a
jpeg, through the same ScreenCaptureKit path the recorder uses. No
Screen Recording permission is involved, so it works on a release build
and inside CI, where the DEBUG-only `screenshot` command does not exist
at all.
The two commands share every decision they could disagree about:
parameter decoding, crop and scale planning, frame capture, output-file
handling and window resolution. `cmux record --region 0,0,420,900` and
`cmux shot --region 0,0,420,900` frame the same rectangle of the same
window, so a detail spotted in a clip can be shot with the same four
numbers.
The image is encoded beside the output path and moved into place, so an
existing file there is replaced only once there is a complete image to
replace it with, and `--out` pointed at a directory comes back as
`invalid_params` rather than deleting it.
Local socket only: `window.screenshot` is not on the `cmux ssh` relay
allowlist, since it reads pixels of windows the remote session does not
own.
## Changelog
Added: `cmux shot` captures a cmux window, or a region of one, to a png
or a jpeg.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Compare recorded pixels through an Equatable value
Optional tuples are not Equatable, so the caption test did not compile.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add a cmux-capture skill for screenshots and clips
`cmux shot` and `cmux record` exist, but an agent looking for a way to show
a reviewer what a change looks like has to find them first. The skill says
which of the two fits a still and which fits a change over time, how to reuse
one `--region` rectangle across both, and what an agent recording its own
cmux needs to know: the recording stops itself, one runs at a time, captions
are drawn in as it goes, and nothing but cmux is in the frames.
`references/commands.md` carries what `--help` cannot: the response fields,
the error codes and the limits the app enforces.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add a capture topic to cmux docs
`cmux docs` is where an agent with no repository checkout looks for a topic,
so the capture skill and its command reference are reachable there too, and
from the agents topic next to the hook integrations.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Guard the capture skill against drift from the code
A doc that names a response field or a state the app no longer emits sends
an agent after something that is not there. Two such errors were in the draft:
it called the recording states `recording` or `stopped` when they are
`recording`, `finished` and `failed`, and called the screenshot `format` field
`png` or `jpg` when the response carries `jpeg` and only the file is named
`.jpg`. Both came from reading the CLI and not the payload.
The guard reads both sides: documented flags against the CLI help, documented
states and formats against their enums, documented response fields against the
keys the payload actually sets, and documented error codes against the ones
the socket methods return. It needs no build, and it runs in the skill
contract workflow when the skill or the code behind it changes.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Probe the capture docs topic from the CLI contract
`test_cli_contract_help.py` runs every probe line against the built CLI, so
the new topic is checked the way the dock topic is.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(skills): complete capture discovery contract
* fix(skills): advertise capture docs in CLI help
* Fix window capture edge cases
Validate and safely clamp region geometry, including origin-anchored crops and finite values beyond integer range. Build GIFs only after their exact frame count is known, staging compressed frames on disk. Keep screenshot output promotion on the waiting socket worker so a timed-out capture cannot later overwrite the requested path.
* fix(capture): preserve directories at promotion
* fix(capture): bound frame work and caption history
* fix(capture): isolate uncooperative frame requests
* fix(capture): bound GIF staging and sampling
* Pin the capture doc claims a reviewer found wrong
Nine doc findings from a review of this skill, and three guards so the same
class of drift is caught next time rather than read past.
The corrections: the error table said `invalid_params` covered an unknown record
subcommand, which never reaches the socket; a region under 8 or over 100000
points, a relative `--out` and an `--out` whose extension contradicts the format
were all undocumented; `not_found` covers an unknown recording and a window that
was never open, not only one that closed mid-capture; `conflict` also answers a
`stop` that raced its start; `fps_effective` needs two frames spanning a nonzero
time, not just two frames.
The guards: the documented region bounds are compared against the `Double`
constants both requests enforce, every relative link in the two pages has to
resolve to a file and to a heading that exists, and every Swift file these
checks read has to be in the workflow's trigger paths. The last one was missing
`WindowRecordingRequest.swift`, so a change to the recorder's own bounds could
not have failed this test.
The link guard immediately caught a forward reference: the skill pointed at
`dogfood-scenarios.md#record-a-clip`, a heading that arrives with the dogfood
record step, so the anchor is dropped until that lands.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Wait on deadlines and make the sample schedule a value
Two CI guards refused the capture branch, and both were pointing at
something worth changing rather than at the gate being strict.
`WindowScreenshotWriteTests` waited by yielding a fixed number of times.
N yields is however long N reschedules take, so the wait tightens exactly
when the runner is busy, which is when a passing test turns into a flake.
It now yields until a five second deadline instead.
`WindowRecordingSampleSchedule` was an enum of static functions that took
the previous target back in as a parameter, so the caller held the state
and the type held none. It is now a struct the sampler owns and steps
forward, which is also why its tests can assert on `targetUptime` after a
served slot rather than only on a returned number.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Pin the writer's error codes, which nothing tests any more
`record start` opens the output file before it answers, so an `--out` the
recorder cannot create fails inside the writer rather than in the request
decoder. `TerminalController.recordingErrorCode(for:)` has no branch for
`WindowRecordingWriterError`, so every one of them falls through to
`internal_error`, and an agent told by the docs to fix its flags on
`invalid_params` reads its own unwritable path as a cmux bug and retries.
This test covered it until the branch was repaired; the case that made it
redundant, `stoppedWhileStarting`, is gone, but the writer errors are not.
It is restored first, failing, and the mapping follows.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Tell a caller when the recorder could not open their clip
`WindowRecordingWriterError.setup` is the file the caller named refusing to
be created, so it answers `invalid_params`; a frame or a close that fails is
the encoder, so those stay `internal_error`.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Pin the limits and codes an agent reads, before fixing them
Three claims the capture docs make that nothing checked.
`invalid_params` tells an agent to edit its flags, so it cannot also mean
"this window has nothing in it": no flag makes an unrendered window
capturable, and the mapper sent all three geometry failures down that
branch. The test now pins the split.
The sample schedule counted one slot past every elapsed slot, so a
sampler arriving exactly on a later slot threw that frame away and waited
a whole interval more for the next one. The new case asks for the slot it
is already in time for.
The limits table was checked for flag names and never for numbers, which
is how it came to promise a gif `--max-width` of 4096 against a ceiling of
1280, and to omit the two gif budgets that have no flag of their own. The
guard now reads the constants and asserts the page states each one, and
the region check compares values instead of mirroring a source line.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Keep an on-time frame, and stop blaming flags for an empty window
`recordingErrorCode` now answers `not_found` for `emptyWindow`, beside
the two geometry failures a parameter can fix. An agent that reads
`invalid_params` changes its flags and tries again, which no flag would
have helped here.
The schedule rounds up to the first slot at or after now instead of
adding one to the elapsed slots, so a target that is exactly `now` is
served rather than skipped, and it steps once more in the one case where
the division rounds down in binary.
Two comments claimed a window resized mid-clip never ends a recording.
That holds for the frame size, which is fixed when the clip opens, but
not for a `--region` the window has shrunk out of: the crop is planned
again per frame, planning throws, and the clip ends in state `failed`
with the frames it had. Both comments now say so.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Document the gif ceilings, and offer an index in --help
The limits table promised a gif `--max-width` of 4096 while a gif stops
at 1280, and said nothing about the two budgets a gif has and an mp4 does
not: 960 frames, and 4000000 pixels a frame. The reference page owns the
limits and `--help` owns the flag list, so the ceilings go here rather
than into a wider flag column.
Also corrected on that page: a region a shrinking window leaves behind
ends a clip, which the region section had not admitted; `conflict` no
longer covers a stop that raced its start, a case that no longer exists;
an `--out` under a read-only volume is `invalid_params`; a window with no
capturable content is `not_found`; and `rpc window.screenshot` takes a
window id, because the CLI is what resolves a ref or an index. The skill
page said `record status` reports the effective frame rate, which only
`--json` carries.
`--help` for both commands listed `--window <id|ref>` while the usage
line and the reference page both offer an index, and the CLI has always
accepted one. Same token in all nine locales, since it is a spec rather
than prose.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(recording): stop asserting how many samples the encoder writes
`twoFramesInTheSameTickStillGetIncreasingTimes` expected a two frame clip
to read back as exactly two samples, and the CI VMs write six: with no
hardware scaler available, the encoder decides how many samples a clip of
two frames a 1/600 second apart ends up with.
The property under test survives. A dropped frame, which is what an equal
presentation time used to cause, leaves a single sample, so a floor of two
still catches it, and the times still have to start at zero, increase, and
put the second sample a tick rather than a frame interval after the first.
The expectations now print the times they read, so the next surprise from
the encoder arrives with its evidence.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(recording): pin that a gif keeps the frames staged before its limit
A gif that reaches maximumStagedBytes mid-clip currently loses every frame:
finish() retries the frame it could not stage, throws again, and never creates
the destination, so the session deletes the working file. That contradicts
fail(_:), which promises "Keep whatever was captured before the failure".
Red before the fix: the finalized gif does not exist.
## Changelog
none
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(recording): pin what a mid-clip crop replan may and may not fail on
A clip's encoded frame size is fixed by its first frame, so replanning a later
frame with the gif pixel ceiling still applied ends a recording over a size
nothing encodes: a window that grows mid-clip fails with gifFrameTooLarge even
though the frame is drawn into the opening size. A region the window shrank out
of is the one resize that still has to end the clip.
Red before the fix: planCrop does not exist.
## Changelog
none
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(recording): keep a gif's staged frames when the last one cannot stage
`finish()` staged the frame it was holding without clearing `pending`, so a
clip that had already reached its staging limit threw again from the one call
that is supposed to finalize it. The session's `fail(_:)` promises the caller
"whatever was captured before the failure", and that promise was being broken
for exactly the recordings that need it.
Drop the one frame that has no room and finalize the frames before it. A clip
with nothing staged still throws, so a recording with no frames still reports
why it has none.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(recording): re-crop a window that grows mid-clip instead of ending the gif
The per-frame replan called the first-frame planner, which applies the gif
pixel ceiling. A window that grew past that ceiling mid-clip therefore ended
the recording over a frame size nothing encodes: the encoded size was fixed by
the first frame, and the grown crop was only ever going to be drawn into it.
`planCrop` plans the crop with no pixel ceiling, and the session uses it for
later frames. The one resize that still ends a clip is a `region` the window
has shrunk out of, because no rectangle is left to sample.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* docs(cli): give the `screenshot` alias its contract row
`check-cli-contract-verbs` reads dispatched verb strings, and the shot arm
dispatches both `shot` and `screenshot`. Only `shot` had a row, so the guard
failed the whole run with "top-level verb screenshot is in no command table".
The first cell may name a verb and its aliases, so both go in the one row.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix: validate screenshot requests and capture docs
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: reproduce answered agent notification ring Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: clear answered agent notification rings Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: clear Codex answered notification rings Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: fence Codex prompt notification retirement Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: normalize agent session fallback matching Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: stamp ordered Codex progress hooks Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: limit Codex progress stamps to safe hooks Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: clear agent rings on terminal input Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: clear agent rings in dock terminals Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: remove unused semantic notification binding Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: isolate agent notification mutation fixtures Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: clear queued mutations between notification fixtures Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: fence superseded agent prompt retirement Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: preserve newer same-session prompts Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: index unread agent prompts Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…15989) * test: reproduce late Agent Chat startup stealing session focus * fix: keep late Agent Chat startup replies from changing selection * Make Agent Chat fork and handoff actions recoverable (#15994) * test: reproduce stuck Agent Chat fork and handoff lifecycle * fix: make Agent Chat fork and handoff actions recoverable * fix(agent-chat): deduplicate reconnecting session actions Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(agent-chat): keep main's pinned tsc in the check script The main merge resolved agent-chat/package.json toward this branch's `bun x tsc --noEmit`, but main had since pinned typescript 5.9.2 as a devDependency and added tests/test_ci_guard_workflow_structure.py, which asserts the exact check script. CI fast guards failed on that assertion. The pinned dependency came through the merge, so ./node_modules/.bin/tsc exists; only the script string was stale. Verified: python3 tests/test_ci_guard_workflow_structure.py passes, and `bun run check` in agent-chat is 106 pass / 0 fail across 18 files. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* Checkpoint terminal scrollback so a crash no longer loses it The 8 s session autosave never captures scrollback, so after a crash or SIGKILL every terminal restored empty; only clean quit, power-off and update relaunch persisted it (#2016, #2194). Add slow, bounded scrollback checkpoints next to the primary snapshot: - at most every 60 s, only after 5 s without typing, driven from the existing autosave timer; - only terminals that produced PTY output since their last capture, flagged by one relaxed atomic load per PTY read in the existing tee; - at most 3 Ghostty VT exports per checkpoint, one per main-queue turn, stopping early on typing or after 50 ms of main-thread capture time; - truncation, encoding and writes on a utility queue, one file per terminal under session-<bundle>-scrollback/, pruned to live panels; - the same eligibility gates as the quit path (running command, hibernated agent), with stale checkpoints deleted. After an unclean exit, startup restore fills each terminal's missing scrollback from its checkpoint; the newer of snapshot and checkpoint wins, and the restore path still applies its own replay gates. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Bound checkpoint capture cost and keep restored scrollback across crashes Review follow-up for the scrollback checkpoints: - Split the capture: only Ghostty's VT export runs on main; reading the export file, CRLF normalization, the 4000-line tail, truncation, encoding and the write run on the utility queue. A terminal whose export alone exceeded the 50 ms budget, or failed, is skipped for 10 minutes. - Seed checkpoints from restored scrollback when a restore completes, keyed by the restored panel ids and without a VT export, so a second crash before the next checkpoint keeps it. - Record scrollbackCapturedAt on snapshots saved with scrollback. A newer scrollback-bearing save wins over a checkpoint even when it deliberately omitted a terminal's scrollback; the 8 s autosave (no marker) still yields to checkpoints. - Re-mark a terminal pending when its export read or file write fails. - Keep a terminal pending for one more checkpoint when output arrived between planning and capture, since the PTY tee runs before Ghostty parses those bytes. - Disable checkpoints under automated test runs, like session restore. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Discard checkpoint exports on quit and seed only after a crash - A checkpoint interrupted by quit or restore now deletes the export files it already wrote synchronously on main (one unlink each) and leaves those terminals pending, instead of handing them to the utility queue, which may not run before exit. - Seed checkpoints from restored scrollback only after an unclean previous launch, not on clean launches or manual reopen, to avoid rewriting up to 400 KB per terminal needlessly. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fix overlapping access in the scrollback checkpoint merge Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Discard checkpoints after a clean exit so a later crash cannot restore them A clean quit saves scrollback into the snapshot but left the checkpoint files behind until the next launch's first checkpoint pruned or rewrote them. Restore reuses snapshot panel ids (workspace and dock terminals), so if that next launch crashed first, its 8 s autosave (no scrollback, no capture marker) matched the old records and the crash restore replayed the earlier launch's scrollback over what the quit saved. Startup now deletes the checkpoint directory when the previous launch exited cleanly, before the restore decision so a skipped restore also drops them. After an unclean exit the checkpoints are that launch's own and are still merged and kept until the restored terminals are seeded. Moves the checkpoint coordinator setup into the AppDelegate checkpoint extension to keep AppDelegate.swift's growth down. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Inject one scrollback checkpoint activity owner and delete removals before returning The PTY tee bridge, the checkpoint coordinator and the persist step now share one `TerminalScrollbackCheckpointActivity` owned by the composition root (`GhosttyApp.terminalScrollbackCheckpointActivity`) instead of a `.shared` singleton, so a caller cannot hand the coordinator a different instance than the one the tee writes. Planned removals (a panel that stopped being eligible) are now applied with `queue.sync` after earlier queued writes, and only captures and pruning go async. A crash before the async block ran used to leave the old checkpoint on disk for the next unclean restore to merge. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Test scrollback checkpoint output interleaving Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Make scrollback checkpoints generation-safe Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Preserve scrollback for idle agent sessions Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Match checkpoint agent liveness policy Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: cover Ctrl-letter key event encoding Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: encode synthetic Ctrl-letter keys Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: compare named Ctrl-letter bytes Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: report named key releases Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: avoid shadowing Ghostty key event type Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: retain synthetic key press capture Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: match Ghostty legacy Ctrl-I and Ctrl-M Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: restore bonsplit pin for macOS compile Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
#15420) * Move browser and tmux OSC filter tests into their package test targets BrowserSystemProxyMirror, BrowserUserAgentPolicy, BrowserNavigationDecisionHandler, BrowserHiddenWebViewDiscardManager and BrowserHTTPBasicAuthPromptCoordinator live in CmuxBrowser, and RemoteTmuxNotificationOSCFilter in CmuxRemoteSession (#13787), but their 68 tests compiled into cmuxTests and ran inside the launched app host. They use only package API, so they move unchanged except for dropping the cmux import. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Drop the blank line the removed import block left Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Keep app shard metadata aligned with package tests * test: give split fixtures deterministic geometry --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add failing tests for agent wake verification The tests need new types, so this commit also adds minimal stubs: AgentWakeVerificationState never changes state, and the Workspace entry points (beginAgentWakeVerification, failAgentWakeVerification, dismissAgentWakeFailure) do nothing. The state machine and Workspace tests fail against these stubs. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Check that a woken agent came back Waking a hibernated agent types its resume command into a fresh shell, but nothing checked whether the agent started. A missing launcher or a gone session left the pane at a shell prompt with no sign of failure. A wake now starts a check for the pane. An agent hook reporting a PID or a lifecycle state counts as success. The check fails when the resume command returns to the prompt first, or when nothing reports within 90 seconds and no live agent process is found. A failure shows a banner on the pane (Retry, Show command with Copy, Close), an "Agent didn't resume" sidebar row, and one notification feed entry. The row is kept out of the saved session. Closing the pane, hibernating it again, or a later agent report clears the check. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Only fail quick resume exits and never retype into a running program A resume command that returns after 20 seconds ran the agent, which the user then quit; for agents without hooks that was reported as a failed wake. It now counts as resumed. Retry is hidden, and refused, while a command still runs in the pane, since the text would go to that program. The failed-wake row survives a sidebar reset, and a retired workspace skips the deadline. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Confirm wakes only from the woken agent's own evidence Hook reports now count only when they come from the woken agent and name their pane explicitly; the focused-pane fallback no longer confirms a wake. A resume command that ran a while is no longer taken as proof: a pending check samples the pane for a live agent process every few seconds instead, and a command that ends before any confirmation fails the wake. The verification deadline constants are nonisolated, which fixes the new Swift warning. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(cloud): cover the reads that find a Cloud machine gone A status read and an access preflight both move the row to `destroyed` themselves when the provider no longer has the machine. Once either does, `destroyVm` can never see the row again (its lookup skips destroyed rows) and the provider-status cron skips it too, so the model-plane revoke and the `vm.destroyed` usage event that the cron performs for the same transition have to happen at these write sites as well. Fails today: both retire the row and record nothing. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(cloud): finish the destroy work where a machine is first found gone Three places retire a Cloud machine's row after the provider says it no longer has the machine: the status read, an access operation's resume preflight, and the stats read. Each wrote `status = destroyed` and stopped there, while the reconcile cron doing the same transition also revoked the machine's model-plane tokens and recorded a `vm.destroyed` ledger event. That difference was permanent, not a race to lose. `destroyed` is terminal: `findUserVm` hides such a row from every destroy request and `reconciliationCandidates` drops it from the cron, so whichever of the three got there first left a machine that never appears as destroyed in the ledger and never has its route tokens marked revoked, with nothing able to finish the job afterwards. All three now go through one `applyObservedProviderStatus` that performs the write and, when the write lands on `destroyed`, the same revoke and ledger event as the cron. The status route hands `getVm` the model-plane revoker it already builds for delete. The new `provider_status_*` destroy reasons join the analytics allowlist so the event says which read noticed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(cloud): derive the observed status inside the shared destroy path Review found that the previous commit took the row's new status from each caller. That let the access preflight and the stats read hardcode "destroyed" for a provider 404, including for a machine with a persistent home volume, which observedDbStatus maps to "paused" because the compute is gone and the machine is not. Those rows were terminalized and billed a vm.destroyed that never happened, and nothing revisits a terminal row to take either back. The status is now derived inside applyObservedProviderStatus, so every entrypoint agrees about what a 404 means. reopenBaseIfProviderDeleted is the one caller that must override it, and passes forceStatus with the reason: its row is a Base's active generation, and leaving it paused would hand the same dead provider id back on every later open. Also threads the model-plane revoker through getVmStats from its route, covers the stats entrypoint including the home-volume case, and fixes the two test-file regressions the previous commit shipped. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * vm: say what a paused 404 row means for credentials The access preflight carried a comment claiming a revoke was pointless there because the credentials were already inert. That is false for the case this branch introduces: authenticateRouteToken and authenticateVmAuthorization both accept `paused`, so route tokens on a volume-backed machine stay valid after this write, for the rest of their 30-day lifetime. Keep the behavior, which is what getVm and the reconcile cron already do for the same observation, and replace the comment with what is true. Also drop the stale "next fleet refresh drops it" line, correct "seven call sites" to eight, stop asserting a resurrection path that is not implemented, and type usageEventSource as VmDestroySource so an unknown source cannot silently degrade in PostHog. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(cloud): reproduce lost destroy ledger retry * fix(cloud): retire missing provider compute durably * fix(cloud): durably drain observed destroy cleanup * fix(cloud): keep destroy cleanup queue fair and indexed * fix(vm): make observed cleanup retries index-safe * test(cloud): reproduce account deletion cleanup loss * fix(cloud): retain pending destroy cleanup on account deletion * fix(cloud): transfer account cleanup to durable outbox * fix(cloud): reject malformed cleanup outbox rows * fix(cloud): validate cleanup outbox shape * Persist explicit destroy cleanup for retry * fix cloud observed cleanup retries and concurrent migration * fix migration retry for invalid concurrent indexes --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat: surface configured actions and launchers * fix(actions): preserve surface button action identity * fix(actions): report exact placements and global edit target * test(actions): pin exact surface placement identity * test(config): preserve surface action references * fix(actions): drop hidden surface action references * test(config): drop references for hidden surface buttons * fix: localize configured actions empty state * fix: sync upstream CmuxFoundation import repair (#13236) Applies upstream 968cef2 to unblock this PR's app-host compilation. * fix(actions): localize discovery metadata labels * fix: localize actions discovery titles and open button Give the Actions menu item and dialog their own whole-string keys instead of borrowing the titlebar-layout debug label and concatenating, and label the open button with the existing Open cmux.json key rather than a raw path. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: document French identity for the actions discovery titles "Actions" is spelled the same in French, and cmux.json is a file name. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat: optionally isolate native agent scratch files * fix(settings): register canonical agent scratch JSON path * fix(settings): document canonical agent scratch in schema * build(schema): regenerate embedded config schema * docs(settings): add canonical scratch key reference * fix: localize canonical scratch schema description * fix: regenerate config schema after current-main refresh
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
App-host test processes on one runner share the app's standard defaults. Every main-window close writes its frame there, and the next process opens its launch window from that frame. A 320-point fixture closed by an earlier process therefore sized the launch window, every createMainWindow() copied it, and split admission (#15392) refused the side-by-side splits the equalize, font-size, remote tmux and manual-unread tests need. A test process now forgets the persisted window frame at launch, next to the existing shortcut resets for the same shared profile. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On macOS 15 runners a titled NSButton has an extra NSTextField descendant, so walking every descendant found 3 text fields for rows with an action button and failed the labels.count == 2 check. The row's labels are its direct subviews; AppKit sizes the button's internal text. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…15921) The blocked-cancellation test held the transport writer only after a fake daemon event arrived, so the writer had to be held within the 1 s attach RPC deadline. When that round trip took longer on a loaded runner, the cancellation write ran first, succeeded, cancelled the write deadline, and the expected transport stop never came (line 182 timed out after 10 s). The test now holds the writer first and then runs the handler a timed-out pty.attach invokes, keeping the writer held until the test ends. The companion test still covers the timeout-to-cancellation wiring. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Austin Wang <austinwang115@gmail.com>
* Add iOS Feed tab with X-style rows and inline agent actions
A new Feed tab, distinct from Notifications, mirrors the Mac's workstream
Feed on the phone as an X-style full-width timeline: agent output renders
inline (preambles, plans, prompts, tool lines) and every respondable
interaction is actionable in place — permission requests (allow once/always/
all/bypass/deny), exit-plan approvals with inline revise feedback, question
option chips with an Other… free-text lane, and turn-complete replies routed
to the owning terminal.
Mac side: a `feed.v1` mobile capability exposes `feed.list` (workstream
items enriched with resolved workspace/surface targets, display titles, and
conversation context) plus `feed.permission.reply`, `feed.question.reply`,
and `feed.exit_plan.reply`, which dispatch to the same handlers the control
socket uses so every entrypoint resolves through FeedCoordinator.deliverReply.
WorkstreamStore gains a monotonic revision with a change hook that emits
revision-only `feed.changed` events to subscribed phones.
iOS side: per-Mac revision-guarded snapshots aggregated across paired Macs,
optimistic local resolution reconciled by refresh, capability-gated targets,
and lifecycle shared with the notification feed. Turn replies reuse
mobile.terminal.input. Strings ship EN+JA in the CmuxMobileShellUI catalog.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Feed round 2: never render wire JSON, keep the Feed decision-dense
Dogfood feedback: raw JSON must never reach the screen, and routine tool
churn buries the rows a user can act on.
Mac: feed.list now parses the ExitPlanMode envelope to plan text (same
parse as the Mac Feed panel) and drops PreToolUse/PostToolUse telemetry
except failed tool results, so the 200-row budget carries notable rows.
iOS: tool payloads humanize to command/file_path/prompt fields or
key: value pairs and never render braces; plan bodies defensively parse
older Macs' JSON envelopes; the view filters residual tool churn from
older hosts; question options render as X-style stacked poll bars (the
custom wrapping Layout is gone) and a pending question always offers the
Other… free-text lane even when its prompt failed to parse.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Fix shadowed normalized() helper in tool-line humanizer
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Lower the keyboard after inline feed replies
Send drops the field's focus, and swiping the feed dismisses the
keyboard interactively, so an abandoned reply never pins it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Feed round 3: toolbar filter menu, reply markers, stable expand, no title
The filter moves from an in-list segmented control to the same toolbar
menu control the Workspaces list uses (Mail-style icon, filled while
narrowing). Replying to a turn-complete row now records the reply on
that row: a quote-reference line names the message being answered and a
"You" bubble shows what was sent, persisted across snapshot refreshes.
Show more expands without animation so the top of the row stays fixed
while the text grows downward. The Feed drops its navigation title.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Restore main's submodule pins clobbered by merge resolution
The conflict-resolution git add recorded the stale local ghostty and
bonsplit checkouts as gitlinks, reverting main's pin bumps (bonsplit
c5452560 carries TabDragTransferRegistration.clearResidualCapability,
which Sources/SessionDragSessionSource.swift requires).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Feed: X-style reply composer sheet; the timeline never hosts a keyboard
Every free-text lane (turn-complete reply, question Other…, plan
Revise…) now opens one compose sheet modeled on X's reply flow: the
quoted message with a thread line, Cancel top-left, a prominent
Reply/Send capsule top-right, auto-focused editor, and swipe-down or
Cancel to dismiss — so the keyboard always has an obvious way down.
Stop rows show a quiet bubble Reply button in place of the old inline
field.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Feed: no divider above the top row
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Feed design pass: prompt rows are authored by You; tighter collapse
Sim UX pass findings: user-prompt rows read backwards ('Claude received
your prompt') — X-style, your prompt is your post, so they now render a
person avatar, 'You' as the author, and 'prompted <agent>' as the
headline. Long outputs collapse at 360 chars / 8 lines.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Feed: unwrap double-encoded plan envelopes instead of rendering them
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Feed: the user's own prompts are not standalone Feed events
Dogfood: the timeline was a wall of 'You prompted …' rows. A prompt is
not news to its author — it already surfaces as the quoted context line
under the agent rows it produced. Filter userPrompt rows on the Mac
(freeing the row budget for notable rows) and on the phone (for older
hosts).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Feed round 4: one action-button family, swipe triage, visible Reply
Dogfood feedback on the controls. Approve/Allow, Revise/Always, and
Deny now share one option-bar-shaped family — accent fill, quiet fill,
and red-tinted fill — replacing the stock bordered mix with its bare
red-on-gray Deny; the overflow ellipsis renders at full strength in a
matching chip. Rows gain mark-read-style leading swipe triage (Done /
Needs Input) backed by a local override that adjusts the badge, filter,
and unread dot without touching the row's answerable state. The stop
row's Reply is an accent-tinted button instead of a quiet footnote, and
the composer's Reply loses its custom capsule so the toolbar's own pill
is the only background.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Fix mobile feed replies and duplicate stop rows
* Use inline reply action styling
* Test Feed reply routing and real terminal submission
* Cover submit-key selection after switching agent providers
* Share Feed prompt submission and honor the active agent
* Require asynchronous prompt submission on the control socket
* Reproduce Enter arriving inside agent paste protection
* Separate agent paste and submit input before acknowledging replies
* Cover provider identity through Feed ingestion
* Preserve Grok identity in Feed events
* fix: remove stale Sentry test target linkage
* Verify Feed submissions against real provider editors
* Improve provider verification harness
* Document Feed session scope and outstanding verification
* Cover refreshing an existing personal development account
* Allow verified personal credential refresh and record connection blocker
* Record tagged dev sign-in gate result
* Record working developer pairing and main refresh scope
* Update readiness fixture for installed bundle evidence
* Cover awaiting notification reply delivery before acknowledgement
* Await terminal submission in notification delivery callers
* Verify notification delivery with refreshed Feed builds
* Keep Feed timestamp formatting on its row model
* Fix main API rename in legacy mobile host
* Adapt feed target resolution to main
* Record main rebuild and queued phone install
* Record transport main refresh
* Record refreshed builds and remaining phone auth blockers
* Test full Feed completion preservation and paged reading
* Open retained Feed messages in a full-text reader
* Exercise Feed reader retry scrolling and dismissal in iOS UI test
* Run focused Feed reading tests through the registered hosted workflow
* docs: explain Feed data flow and record delivery and new scope
* test: require Feed-scoped search and conditional inline expansion
* fix: inline Feed expansion, scoped search, shared controls and event navigation
* fix: isolate Feed search navigation and apply exact computer scope
* fix: contain inline touch targets and seed navigation tests through store state
* fix: handle Feed search in the workspace preview scaffold
* Document Feed navigation verification and queued phone update
* Test unified Feed open action
* Use one Open action for Feed destinations
* Persist feed replies and mirror notifications
* Keep empty notification feed revisions compatible
* Document durable notification feed inputs
* Preserve durable feed order on replay
* Cover exact feed event reply identity
* Record unified feed open action
* Persist feed generation across restarts
* docs: point feed data flow to current source
* docs: record current feed build status
* Add swipeable multi-question Feed answers
* Render Feed event content as Markdown
* Fix Markdown intent raw value conversion
* Fix Markdown sheet navigation modifiers
* Test Markdown preview truncation and expansion offsets
* Preserve rendered Markdown while truncating Feed previews
* Keep question answer composition on its draft type
* Test block Markdown rendering in Feed previews
* Render block Markdown in Feed row previews
* Record Feed Markdown verification and paired build delivery
* Test Feed first row spacing beneath the toolbar
* Remove the empty large-title area above Feed rows
* Use shallow checkout for manually dispatched iOS checks
* Record Feed spacing proof and authenticated phone delivery
* Keep current Feed delivery status consistent in scope notes
* Quiet the Feed Reply action to a text-scale secondary control
The filled blue arrow and 44-point layout frame made Reply louder than
the row content and added a gap above it. Render an outline icon and
footnote label in secondary color, keep blue for links, and extend only
the hit area to 44 points so row layout stays compact.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Record the quieter Feed Reply action in scope notes
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add iMessage-style Feed quote bubbles behind a CMUX Labs switch
Quoted prompts render as outlined bubbles with a leading tail, and a
recorded reply shows the answered message as an outlined quote above a
filled reply bubble, matching iMessage inline replies. CMUX Labs >
Feed Bubble Quotes switches between this and the original leading-bar
quotes; it defaults on in DEBUG and is unavailable in production.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Use the Messages tail geometry for Feed quote bubbles
The first tail flattened the bottom-leading corner into a slab. Draw the
body inset by the tail width and hook the tail out and back into the
bottom edge, as Messages does, so stroked and filled bubbles read as
iMessage bubbles.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Record Feed bubble quotes Labs switch in scope notes
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Point Feed bubble quotes of the user's words to the right
A quoted prompt is the user's own message, so draw it as a sent bubble:
trailing-aligned, accent-tinted, with the tail mirrored to the right. In
a recorded reply the answered agent text stays a gray leading quote and
the user's reply becomes a filled accent bubble on the trailing side.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Smooth the Feed bubble tail join into the bottom edge
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Record right-pointing Feed quote bubbles in scope notes
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add failing test for tapping links in Feed previews
Feed previews render Markdown links but their text view ignores touches,
so tapping a link such as a PR number does nothing.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Open Markdown links tapped in Feed previews
The preview text view is non-interactive so row gestures keep working,
which also swallowed link taps. A tap recognizer on the container now
begins only when the touch lands on rendered link text, resolves the
destination from the laid-out glyph, and opens it through SwiftUI's
openURL like the Feed's other Markdown text. Other taps pass through.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Show each Feed event's workspace, with an optional tab setting
Finished-turn rows drop the "finished a turn" verb so the author line
reads agent · workspace. Every row shows the workspace it came from
(its title, else the cwd folder). Settings > Display > Show Tab in Feed
appends the event's tab for people who organize agents by tab.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Open Feed rows inside the Feed tab like Notifications rows
Tapping a Feed row now opens the event's workspace or tab, and the
workspace pushes onto the Feed tab's own navigation stack (new
agentFeed deep-link origin), so Back returns to the Feed instead of
jumping to Workspaces. Buttons, links, and See more keep their own
taps; VoiceOver gets an Open action, and rows get stable identifiers.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add Feed offline banner and a Retry action
When the Feed cannot refresh but has cached rows, a banner above them
says it is offline or that a Mac needs an update, matching the
Notifications tab, so stale activity is not mistaken for live. The
full-screen unavailable state gains a Try Again button.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add a CMUX Labs switch to let the Feed replace Notifications
Feed Replaces Notifications hides the Notifications tab on iPhone and
its destination in the iPad sidebar, moving a selected Notifications
tab to the Feed, so the replacement can be dogfooded before the tab is
removed. DEBUG-only; production keeps the Notifications tab.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Pin Xcode and route forks in the Feed reply provider workflow
Both macOS jobs now select Xcode through scripts/select-ci-xcode.sh and
fall back to a GitHub-hosted runner outside manaflow-ai, satisfying the
pinned-Xcode and fork runner routing guards.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Show failed Feed replies and keep the text for retry
A terminal reply that fails now leaves a red note on its row instead of
only a log line. When the owning Mac was offline the note says the
reply was not sent; when the request reached the Mac or its fate is
unknown it says delivery could not be confirmed and offers Open
Terminal, so the user checks before retrying and the text is never
typed twice unseen. Try Again reopens the composer with the text.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Order Workstream log appends and let callers await them
Each append was an independent Task, so two mutations of one item could
reach the JSONL log out of order and a restart would replay the older
version. Appends now form one ordered chain, and flushPersistence()
awaits it. The restart tests use it instead of sleeping 50 ms, which
the test-determinism guard rejects.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Record the overnight Feed readiness pass in scope notes
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Record Feed usage events like the Notifications tab does
Appends agent Feed app events (opened, closed, filter changed, item
opened, reply succeeded, reply failed with its delivery state) to the
diagnostic taxonomy and records them from the Feed store view, so
replacing the Notifications tab does not lose usage visibility. Events
carry only counts, never content.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add a UI test that a Feed row tap opens its destination
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Record Feed analytics and local test results in scope notes
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Record overnight CI findings and the Feed UI test lane gap
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Fix question pager clipping, nav buttons, and duplicate notification rows
A paged TabView takes its proposed height and centers overflow, so tall
question pages clipped mid-glyph at 212 points. The pager now measures
each page and sizes the frame to the current page. Previous and Next
join the row's shared action-button family (quiet fill and accent fill,
trailing chevron) and navigation is never disabled; only Submit gates
on complete answers. The Feed's actionable banner records its history
notification under the feed event's id, so the mobile Feed's existing
id dedupe drops the duplicate Notification row beside the real one.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Keep agent-hook notification echoes out of the mobile Feed
Agent hooks deliver a notification banner beside every feed event, and
the mobile Feed's history merge re-listed those banners as their own
rows, so the phone showed a Notification row duplicating the question,
permission, plan, or turn-completion row beside it. Notifications now
record whether an agent context produced them (threaded through history
records and session restore, optional so old data still decodes), and
the mobile Feed merge skips agent-produced records; human-facing
notifications such as cmux notify and watchers keep their rows.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Record F46 pager fix and F47 echo suppression in scope notes
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Extend the session notification snapshot init with isAgentEvent
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Order snapshot init arguments to match the declaration
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Record overnight verification results and new gaps in scope notes
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Exclude unflagged notification history from the mobile Feed merge
Records written before the isAgentEvent flag existed are mostly agent
echoes, so requiring an explicit human-facing flag removes the stored
Notification rows the user still saw; every notification remains in the
Notifications tab.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Show only the user's reply bubble, matched to quote size, with a soft tail
The reply marker repeated the agent's message as a gray bubble right
under the same text in the row; bubbles now belong to user text alone,
so the marker is just the reply, at the quoted-prompt bubble's footnote
size and padding. The bubble tail ends in a lifted rounded tip instead
of a sharp point. Scope records G22: codex emits one turn through two
lanes with different reason texts, so the equal-reason stop dedupe
keeps both rows.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Unread-by-default Feed rows and one row per provider turn
Every new event now arrives needing the user (dot, badge, and Needs
Input filter) until a visible, explicit interaction clears it: an
answer, a reply, opening the row, reading full text, or the swipe
triage. Scrolling past marks nothing. Read state and its baseline
persist phone-locally. Stops that describe the same turn through two
provider lanes (a truncated single-line preview beside the full text)
collapse onto one row keeping the fuller text within a two-minute
window, fixing the duplicate codex response rows.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Trim the read-row set with a prefix rebuild
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Hide the Notifications tab by default; Feed replaces it
The replacement ships for everyone: the Notifications tab and its iPad
sidebar destination stay hidden unless Settings > Display > Legacy
Notifications Tab brings them back, keeping the legacy screen hidden
but accessible. The former DEBUG-only Labs switch becomes this shipping
setting with its default flipped.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Expand the replied-to event to its full text in the composer
The reply sheet focuses one event, so its quote now renders Markdown
and offers See more when truncated: the full text loads through the
same paged reader the full-text view uses and the sheet scrolls it in
place, with a retry label on failure.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Truncate Feed previews with one layout pass instead of a binary search
Scrolling re-measures rows constantly, and each truncated row paid a
binary search of full TextKit layouts to find its cut point. A single
layout pass now reads where the visible lines end, then steps back
through composed characters until the See more suffix fits, removing
the repeated relayouts behind the reported scroll hitches.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Fix review findings in the unread, dedupe, and notification-flag paths
The feedReplacesNotifications key moves out of the DEBUG block so
Release builds compile. Read state for stops is also keyed by turn, so
the same-turn collapse cannot resurrect a read row when the fuller
lane replaces it, and read keys evict oldest-first through an ordered
list instead of arbitrary Set order. Same-turn matching now requires
the shorter reason to carry the truncation ellipsis, so two distinct
complete reasons never merge. A failed feed fetch leaves its Mac
pending so an equal re-emitted revision still refetches. The truncated
preview gets a maximumNumberOfLines backstop, and notification
snapshots with unknown provenance restore as agent-produced so legacy
banners cannot re-enter the mobile Feed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Fix mobile feed full text layout
* Fix mobile feed reply expansion and empty rows
* Fix feed UI test scope
* Fix async feed notification tests
* Polish mobile feed decision controls
* Fix mobile feed rendering and scroll performance
* Add mobile feed interaction coverage
* test(ios): keep feed UI test build actor-safe
* fix(ios): use supported feed control sizing
* test(ios): expose paged feed questions as containers
* test(ios): swipe the feed pager directly
* test(ios): cover full text inside feed replies
* test(ios): make full reply fixture addressable
* fix(ios): order full reply fixture arguments
* test(ios): verify feed pager intrinsic layout
* fix(ios): measure paged feed questions after sizing
* fix(ios): top align paged feed questions
* Make feed link hit testing layout-safe
* Fix CI cancellation guard assertion
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: austinywang <austinwang115@gmail.com>
Co-authored-by: Lawrence Chen <54008264+lawrencecchen@users.noreply.github.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Agent TUIs print key hints all over their output:
ctrl+o to expandunder a collapsed tool result,esc to interruptin the status line,shift+tab to cyclebeside the mode. With the newagentActions.keyHintssetting on (off by default, Settings > Terminal > Clickable Agent Key Hints), clicking one of those hints in a terminal pane presses its keys. Hovering a hint underlines it, shows a pointing hand, and a tooltip such as "Click to press ⌃O (expand)".How it works:
AgentKeyHintDetector(CmuxTerminalCore) reads the clicked viewport row on its own (no joining of soft-wrapped lines): one or more chords (ctrl+x,alt+x,shift+tab,esc,tab,enter,space, arrows,⌃O-style glyphs,ctrl+x ctrl+ksequences), optionallyagain, thento/forand an action that ends at·,|, parentheses, a two-space column gap or the end of the line. OpenCode's bareesc interruptform matches only for its known actions. Bare key words (up to date,end to end) count only before a hint verb, andreturnis not a key word. Keys that act on whatever UI is live (Enter, Esc, Tab, arrows, bare letters, Home, End, page keys) count only in the agent's live region: the viewport is at the bottom and the row is no more than 8 rows above the cursor row, or below it, where the status line, input box, footer and permission prompts are drawn. Ctrl and Alt chords (ctrl+o to expand) count on any row, so prose such as "Scroll down to view the full log" stays inert.ctrl+c,ctrl+d,ctrl+z,ctrl+\) are never clickable, and are never sent even if a keybindings file maps a hinted action to one. Any chord with Ctrl andc,d,zor\counts, whatever Shift or Alt is added. A drag, a selection or a Shift/Option/Control click presses nothing. A click presses its hint once the system double-click interval passes without a second press, so a double or triple click selects text and presses nothing.~/.claude/keybindings.jsonbinds that action to a different one. The file is checked on click and cached by modification date; hover reuses a check for up to 5 seconds. Codex and OpenCode render their hints from their own keymaps, so their printed chord is sent. Then, if the pane's Ghostty keybindings would consume the chord (ghostty_surface_key_is_binding), the key's legacy encoding is sent instead (ctrl+oas0x0f,shift+tabasESC [ Z, arrows asESC [ A...), so the terminal binding doesn't run; a key with no legacy encoding is skipped.OS-level remaps (Karabiner, modifier keys in System Settings,
hidutil) and a tooltip that names the physical key to press come in a follow-up.Part of #15223. Stacked on #15234.
Testing
Ran locally:
swift test --package-path Packages/macOS/CmuxTerminalCore --filter 'AgentKeyHint|TerminalLegacyKeyEncoding': 29 tests in 6 suites passed (detector incl. live region and prose, signal chords with extra modifiers, chord resolver against keybindings files, mtime cache and hover max age, live region rule, deferred press timing, click policy, legacy encodings).swift test --package-path Packages/macOS/CmuxTerminal --filter TerminalTextRegion: passed (a viewport row selects exactly that row; a blank row below short content reads blank; a soft-wrapped line reads one row at a time).swift test --package-path Packages/macOS/CmuxSettings --filter SettingCatalogTests: passed, including the new default-off test.python3 tests/test_cmux_schema_parity.pyandpython3 tests/test_cmux_settings_supported_paths.py: passed.python3 scripts/verify-local.py: 15/15 checks passed.python3 scripts/localization_catalog.py check: 0 parity errors.Not run locally: the app compile and
cmuxTests. CI builds the app and runs the newcmuxTests/TerminalAgentKeyHintTests.swift, which checks that a hint click in an agent pane enqueues one key (two forctrl+b ctrl+b), that nothing is clickable with the setting off, that bare keys are clickable only in the live region, that a running agent wins over a stale lifecycle, without an agent lifecycle, on prose orctrl+clines, and that a mouse-capturing agent needs a Command-click. The mouse release, hover underline, tooltip and Ghostty binding fallback are not yet exercised in a live build.Localization: new Settings title and subtitle in the CmuxSettingsUI catalog, and the Settings search title and alias plus both tooltips in the app catalog, all in en, de, fr, ar, es, zh-Hant, zh-Hans, ko and ja.
Changelog
Added: Key hints that Claude Code, Codex and OpenCode print in a terminal, like "ctrl+o to expand", can be clicked to press those keys (Settings > Terminal > Clickable Agent Key Hints)
Demo Video
Pending dogfood.
Checklist
🤖 Generated with Claude Code
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.Summary by cubic
Makes the key hints Claude Code, Codex, and OpenCode print in terminal panes (like
ctrl+o to expandoresc to interrupt) clickable, behind a new off-by-default setting at Settings > Terminal > Clickable Agent Key Hints. Clicking a hint presses its keys; hovering underlines it and shows a pointing hand and a tooltip. The exact/goal resumeprompt Codex prints is also clickable, sending the command and Enter through the same deferred hint path.How clicks behave
ctrl+c,ctrl+d,ctrl+z,ctrl+\) are never clickable or sent, even when a keybindings file maps an action to one, including with Shift or Alt added./goal resumeprompt clicks only in the live row region with an allowlisted command; no other shell text is clickable.~/.claude/keybindings.json, read on click and cached by modification date.main; only the agent key hint changes belong to this PR.Written for commit b231717. Summary will update on new commits.
Migrated from #15305 after correcting the PR author identity. The head branch and commit history are preserved.