Show the cmd-hover affordance over GitHub references - #15259
teamleaderleo wants to merge 232 commits into
Conversation
Stack invitations carry a stored role applied on accept; reusable member-only links redeem through an atomic claim; a per-team advisory lock keeps one admin. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Checkout, portal, subscription, and plan take an explicit teamId and require team_admin. Requests without teamId keep the legacy implicit team. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Profile, emails and sign-in methods, notifications, sessions, API keys, and account deletion now live under /dashboard/settings; /dashboard/team redirects old hash links. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ance Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Cmd-hovering a file path turns the cursor into a pointing hand, so you know before you click that there is something there. Cmd-hovering `manaflow-ai#15221` or a commit SHA gave you nothing: the reference was clickable but invisible, which is the worst of both, since you only find out by clicking and seeing whether a browser opens. The hard part is that a bare `manaflow-ai#123` only resolves against the pane's repository, and pointer motion cannot await an actor. So the cache grows a synchronous read: `cachedSlug(forDirectory:)` reports what a finished lookup left behind, and says `unresolved` rather than "no repository" when nothing is known yet. Confusing those two would mean a hover arriving before the first lookup decides there is nothing to click, which would then be memoized. An unresolved directory starts one background lookup and re-runs the hover when it lands, so the affordance appears under a pointer that never moved. The answer is memoized per terminal cell, because reading the visible line costs a terminal text read and pointer motion inside one cell cannot change it. `pointerCell(at:)` is that geometry split out of the snapshot so callers that only need to know whether the cell changed do not pay for the read. A path still wins where there is one, so a file named `#1` opens as a file. Only text that resolves to nothing on disk is offered to GitHub, matching what cmd-click already does. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Repository: manaflow-ai/cmux/.coderabbit.yaml Review profile: ASSERTIVE Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
📝 WalkthroughWalkthroughThe terminal detects GitHub issue, pull-request, and commit references. It uses explicit repository slugs or resolves a slug from the pane’s repository. Eligible references open through terminal link routing. ChangesGitHub references in terminal
Priority: ⬇️ Low Estimated code review effort: 3 (Moderate) | ~25 minutes Change: Feature Sequence Diagram(s)sequenceDiagram
participant User
participant GhosttyTerminalView
participant TerminalGitHubReferenceClickPolicy
participant TerminalGitHubReferenceDetector
participant GitHubRepositorySlugCache
participant TerminalLinkOpenCoordinator
User->>GhosttyTerminalView: Command-click a terminal token
GhosttyTerminalView->>TerminalGitHubReferenceClickPolicy: Evaluate runtime outcome and visible line
TerminalGitHubReferenceClickPolicy->>TerminalGitHubReferenceDetector: Detect reference and repository requirement
TerminalGitHubReferenceDetector-->>TerminalGitHubReferenceClickPolicy: Return reference or repository-resolution decision
GhosttyTerminalView->>GitHubRepositorySlugCache: Resolve pane repository slug when needed
GitHubRepositorySlugCache-->>GhosttyTerminalView: Return cached or discovered slug
GhosttyTerminalView->>TerminalLinkOpenCoordinator: Open resolved GitHub URL
Suggested reviewers: Merge Risk: 🔵 Low · up to The change is mergeable with owner awareness: scrolling or new output can leave the hover cursor stale, overlapping lookups can schedule redundant tasks, and the hover test may fail under slow CI conditions. Security Architecture ReviewSecurity architecture risk: 🔵 Low · up to Opening a reference still requires a user click, and generated links start at GitHub. No security bypass was established in the inspected flow. Repository freshness and the full scope of the stacked changes remain incompletely verified. Retained concerns Security review detailsSecurity Blast Radius
Trust Boundaries and Controls
Resilience and Maintainability Implications
Important Pre-merge checks failedPlease resolve all errors before merging. Addressing warnings is optional. ❌ Failed checks (3 errors, 1 inconclusive)
✅ Passed checks (21 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 43.82% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 89 functions across 8 files. (3 skipped: 1 unsupported, 2 too large.) Full details: Cmux Swift ConcurrencyExplanation The diff adds unowned fire-and-forget tasks in Resolution Store the hover lookup task on Full details: Cmux Architecture RethinkExplanation The new Resolution Make one repository-slug owner authoritative. Remove the lock-backed Full details: Cmux No Test Or Debug Seam In Production SourceExplanation The PR adds test-only observability to production Swift. Resolution Remove the GitHub fields from ✨ Finishing Touches 💡 1🧪 Generate unit tests (beta)
🛠️ Fix failing CI checks 💡
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
All contributors have signed the CLA ✍️ ✅ |
Review of the hover commit found the memo answering for a cell the pointer had already left. `||` short circuits, so a cell that resolves to a file path never reaches the GitHub branch, and the memo is a single slot invalidated only by being overwritten. Cross a path-bearing cell, let the terminal scroll, come back, and the cursor uses the answer from a line that is no longer there: no pointing hand over a reference that a cmd-click would open, or a pointing hand over prose. Clearing the memo when a path wins keeps the saving the short circuit exists for. The background repository lookup had the same class of problem from the other side. Its completion is the only asynchronous way back into the hover, and it checked the modifier state and nothing else. `updateWordPathHover` falls back to the last in-bounds point it saw, so a lookup landing after the pointer left the view, or after the pane was closed, would push a pointing hand with nothing left to pop it. It also dropped the selection check every synchronous caller makes, so a lookup landing mid cmd-drag lit the affordance over a selection. The completion now proves it is still a live hover before recomputing. The memo's own comment claimed the next pointer move corrects staleness. Motion inside a single cell does not, so the comment says what the code does. The hover debug seam now also reports the GitHub answer and the cell it was computed for, which is what a tour needs to walk path cell to reference cell and back and see the memo follow. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Review of Review
Explicitly attacked and found sound: the synchronous mirror (both write sites are paired, the generation guard covers Fixed 1, 2, 3, 4, 5. The memo is cleared when a path wins, so the short circuit keeps its saving without keeping a stale answer. The lookup completion now proves it is still a live hover (cmd held, pointer still resolvable inside this view) and passes Partly 6: Left The view-level regression test itself. The seam is in place but the test is not written: the feature's async half gates on hardware modifier state that The three low findings in 7. The wasted lookup round is one task and one actor hop, bounded. The pane-identity asymmetry costs at worst a lookup against a directory nobody is hovering. The extra text read is once per cell change, against a path hover that pays the same cost on every event. No local build: |
Adds the typed route tree, session endpoint, flat search params, and the shell, account menu, and team scope ported off next/navigation. Section routes are placeholders until each section is ported. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Billing, TestFlight, CodeRouter, Cloud, Mobile devices, Settings, Teams,
Vault, home, and legacy redirects now run as typed TanStack routes with
TanStack Query data. Server-rendered reads move behind JSON endpoints:
/api/dashboard/{billing,coderouter}, /api/teams/[teamId]/billing,
/api/testflight (GET), /api/vm/access-grants, /api/vault/summary,
/api/vault/cli/auth/client, and /api/vault/sessions/[id]/head.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Next re-asserts the URL from a useInsertionEffect after every external history write; TanStack's history treated that echo as a navigation and scheduled a React update inside the insertion effect. Next's __NA writes now skip TanStack's subscribers, and the history patch is removed when the SPA unmounts. CodeRouter writes use useMutation, route titles come from typed staticData, and new strings are translated for every locale. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The debug hover payload already reported the GitHub half of the answer, but nothing read it. A scripted tour cannot: its hover step holds no modifiers, and this affordance only exists while Command is down. The harness does hold it, so assert the three cases directly. The slug-qualified form has to light on the first hover because it needs no lookup. The bare form is allowed to take a hover or two, since the first one is what starts reading the pane's remote, so that test polls. The ordinary word is checked repeatedly rather than once, because the failure worth catching is a late lookup turning a plain word into a link after the fact. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review points, same class as the ones on the click tests. The ordinary-word hover test asserted only that the affordance stayed off, which is also true when the pointer missed the token, when the surface is gone, or when the feature is deleted. Each hover now also asserts the harness reported `gitHubHoverCell`, which is set the moment the pointer resolves to a cell and before any decision about what is in it. Both fixtures claimed `manaflow-ai/cmux`, the repository CI runs from, so resolving to the checkout instead of the fixture directory would have looked the same. They now claim `manaflow-ai/cmux-reference-fixture`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review: cmd-hover test coverage (
|
CI failure attributionCI failed on
Not re-run automatically: Written by |
|
Not mergeable right now, and it is not this branch.
That guard is Everything else red here cascades from it: the Linux gate declines macOS, so #15360 ("dogfood: give the agent activity reorder tour paths globs") already |
This branch was about to revert the tour fix, and now it does notChecking why CI was red here, I compared this branch against #15221 and found a The tour scenario here was the pre-rewrite copy. Both PRs are based on The cause is that the two earlier merges of Fixed by merging the current New head: On the diff you are readingThis PR is not stacked on #15221 in GitHub's eyes, because #15221's head branch Once #15221 lands, the rest drops out on its own. The red checks were not this branch
The diff touches no So: infrastructure flake, re-running. I am not claiming this branch is green |
Correction: half of that red was this branch, and I said it was notAn hour ago I wrote here that the guard failures "were not this branch" and
So the repository already had a guard for exactly the reversion I described in What I did wrong is specific: I ran the two suites I had guessed were involved
The re-run is on One thing worth keeping: the near-miss in the comment above needs no new guard. |
Out-of-band UI test run at
|
* test: cover descriptor-backed artifact thumbnails * fix: decode artifact thumbnails from verified descriptors --------- Co-authored-by: Leo Li <cheerleaderleo@outlook.com>
…s at once (manaflow-ai#15869) * Revoke a removed member's team VM access on the Stack webhook Stack team_membership.deleted now detaches every tunnel of that user from the team network and drops their identity snapshot, so open terminals and browsers lose the route at once instead of after the 10 minute cron. Tunnel enrollment reconciles against the complete, fresh team list and never detaches when that list is incomplete, so a Mac keeps every team network it belongs to. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Keep Cloud surfaces from several teams open across team switches Each Cloud provider, workspace binding and browser panel records its owning team, and every VM request for a live surface sends that team instead of the active one. A team switch now only rescopes the sidebar and new machines; open terminals and browsers of other teams stay connected, including after restore. A 404 vm_not_found, 403 or vm_owner_mismatch is permanent: the pane shows that access was lost instead of freezing on its last frame. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: temporary hang diagnostics for Cloud graph suites Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
) * Add dictation accessibility regression tests Cover what third-party dictation tools need from the terminal's accessibility element: AXValue shows the screen so a tool can confirm its insertion (manaflow-ai#722), setting AXSelectedText types at the cursor (manaflow-ai#4953), a multi-line value arrives as one bracketed paste, and an unbound Command chord posted by another process does not type into the terminal (manaflow-ai#4153). These fail before the fix that follows. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Let third-party dictation tools insert into terminals Dictation tools that insert through accessibility could not confirm their insertion, could not use AXSelectedText, and a multi-line dictation ran each line as its own command. - AXValue now returns the active screen from a 500 ms snapshot, with NumberOfCharacters, VisibleCharacterRange, Line(for:) and String(for:) getters, and posts valueChanged after an AX insertion (manaflow-ai#722). - setAccessibilitySelectedText commits at the cursor, and both setters are reported as settable (manaflow-ai#4953). A client that writes back the value it read with its text spliced in gets only its text typed, not the whole screen. - Text with an interior line break goes through the paste path, so a bracketed-paste-aware shell or agent gets one block. Single-line text, including a trailing newline sent as Return, keeps typed semantics. - A Command chord that misses the menu and every Ghostty binding is dropped when another process posted it, so a dictation hotkey such as Cmd+Option+C no longer types a stray character (manaflow-ai#4153). The accessibility overrides move to GhosttyNSView+Accessibility.swift. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Harden the dictation AX splice and narrow the foreign chord drop Review follow-ups: - Diff an AXValue write against the last 8 distinct values handed to AX clients, so a tool that writes back an older read after the screen changed still gets only its text typed. Short values need a pure insertion, so a literal isn't trimmed to match a two-character prompt. - Take the multi-line paste path only on a live surface, where the paste and the trailing Return share one input sequence, and end any IME preedit first. - Drop only Command+Option chords posted by another process, the push-to-talk hotkey case, so remote-control and automation tools keep their other unbound Command chords. - Cover a stale read, a multi-line value with a trailing newline, and assert the keyboard chord in the foreign-chord test is local. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: cover delayed accessibility edits Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: retain accessibility reads for delayed edits Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: cover accessibility history retention Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: an exact older AX read wins over a partial newer match Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: prefer an exact AX read over a partial newer match With 30 s of history, a write spliced into an older read could match a newer screen that shares most of its text and paste the older screen's differing tail. Look for a pure insertion across the whole history first. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: announce AX inserts after terminal screen updates --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…anaflow-ai#10277) * perf: read listening ports from the kernel instead of spawning lsof The batched port scan ran `lsof -nP -a -p <pids> -iTCP -sTCP:LISTEN -Fpn` every two seconds. Before lsof answers it calls close() on every descriptor number up to `kern.maxfilesperproc`, which is 138,240 on current macOS, and it forks a child that repeats the sweep. That is far more work than the lookup itself, and it repeats for the life of the app. `ListeningPortLookup` asks the kernel directly through libproc (PROC_PIDLISTFDS + PROC_PIDFDSOCKETINFO) and keeps only TCP sockets in TSI_S_LISTEN. Measured against lsof over 483 PIDs the results are identical, and one scan drops from 200ms to under 7ms. Completeness semantics are unchanged: a PID we may not read is a miss only when its identity is also unreadable while it is still present, so a panel behind the root `login` process can still retire its ports. * perf: stop local port scans when the ports detail is hidden Remote port scanning already follows `sidebar.showPorts` (issue manaflow-ai#6123), but the local scanner kept running its two-second sweep no matter what. Nothing displays or reports the result while the detail is hidden, so the work is pure cost. `SidebarWorkspaceDetailDefaults.portScanningEnabled` now holds the one rule both paths read, and `TabManager` pushes changes to the local scanner alongside the remote sessions it already updates. Note this also blanks `listeningPorts` for local workspaces in the control socket and CLI summaries while ports are hidden, which is the same trade remote workspaces already make. * Stop and restart the local port scan cleanly with the ports detail Hiding the ports detail now also invalidates a burst that is already running, the same way unregistering the last panel does, so its remaining timers never scan. Showing it again rescans every registered panel once, since ports that opened or closed while hidden were never seen and an idle panel would otherwise keep a stale list until its next command. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Address review: clear hidden ports, move publication tests off lsof - Hiding the ports detail now publishes no ports for every tracked panel and agent workspace, so the socket, CLI, custom sidebars and the command palette (which read published ports without the sidebar's visibility gate) don't keep a list frozen at the moment scanning stopped. A fresh agent request ID fences off any scan still in flight. - A queued follow-up panel or agent scan no longer runs after scanning is turned off; a dropped agent request clears its in-flight mark. - PortScannerPublicationTests fed ports through a fake lsof, which the scanner no longer calls. Both runners now supply ports through listeningPortsProvider, and the churn test counts per-PID lookups instead of lsof argv lists. - Test for the cleared list; normalize the pbxproj; drop stale lsof wording from two doc comments. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Pass sets to the snapshot reconcilers when clearing hidden ports Fixes the CI compile error in the previous commit: remove(keys:) takes a Set, not an Array. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Leo Li <cheerleaderleo@outlook.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(cloud): cover idle hover accessory hit testing * fix(cloud): keep idle row clicks available * fix(cloud): hide stale row hover controls * fix(cloud): keep unavailable display actions inert * test(cloud): validate idle row hit target coordinates
* test: cover workspace group anchor behavior * fix: route workspace group headers safely * test: preserve generated group header fixtures * fix: preserve generated group anchor state * Update group anchor lifecycle fixtures * Align group reorder fixtures with anchor semantics * Fix remaining group fixture compile error * Preserve generated-anchor reorder fixture coverage * cmux: follow current Bonsplit tab ID API * fix: close generated anchors consistently * fix: scope generated anchor cleanup * fix: protect confirmed groups during deletion * ci: declare the vendor/bonsplit bump as forward submodule-forward-only: allow vendor/bonsplit The Bonsplit update is a forward move whose merge ancestry is hidden by the shallow submodule checkout. * ci: preserve Bonsplit forward declaration after merge submodule-forward-only: allow vendor/bonsplit * test: cover confirmed group deletion cleanup submodule-forward-only: allow vendor/bonsplit Use the Bonsplit main commit that contains the existing tab ID API and current Liquid Glass changes. * fix: preserve generated anchors with Dock content * fix: retain current Bonsplit API compatibility
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…16150) * ci: report a cancelled run instead of a routing failure The `tests` gate in ci.yml and the `ios-tests` gate in test-ios.yml both run with `if: ${{ always() }}` so that they always report, because both are in the branch protection required list. A cancelled run therefore reaches them with every job it depends on reporting `cancelled`, and their routing check fails with only `changes: cancelled` or `detect-ios-changes: cancelled` to show for it. That message describes a routing defect, and the pull request does not have one. It cost three pull requests a diagnosis each this week before the answer turned out to be the same every time: rerun the cancelled run. Reruns of two of them came back green with no change to the branch. So say that instead. Each gate now has a step that runs only when the run was cancelled and prints what happened and what clears it. Both gates still fail on a cancelled run. A cancelled run tested nothing, and reporting success would let a pull request merge on a green required check that no suite stood behind. `cancelled()` is true only for a cancelled run, so a job that timed out still reaches the routing check and fails there, which is where a timeout belongs. `cancelled()` is available in `jobs.<job_id>.if` and `jobs.<job_id>.steps.if` only, not in a step's `env:`, so this is two mutually exclusive steps rather than one step reading a variable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ci: detect the cancellation in the gate script, not in a step condition The first version of this used a second step gated on `cancelled()`. That only covers a cancelled workflow run: `cancelled()` says nothing about an individual dependency, so a job cancelled on its own still fell through to the routing verdict and still reported as a routing defect. Both gates already read every dependency's result out of `toJSON(needs)`, so the cancellation is visible there. Checking it in the script covers a cancelled run and a cancelled dependency with one branch, needs no step condition, and drops the second step. The verdict is unchanged in both other cases: a run where everything reported still exits 0, and a genuine routing failure still exits 1 with the message it had before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…i#14868) * Add setting actions, setting presets, and cmux config set cmux.json actions can now change settings: "type": "setting" with a path and one of set, toggle, cycle, or unset, and "type": "settingPreset" to apply a named partial settings object from the new top-level settingPresets. `cmux config get|set|unset|toggle|cycle|preset` drive the same JSONConfigStore.apply path, so every entrypoint validates against the schema, keeps comments and unrelated keys, and publishes one atomic write under the cooperative writer lock. Setting actions only run when the global cmux.json declares them; project configs and packs can't ship a button that rewrites global settings. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Address review: empty preset objects, fail-closed trust, docs - A preset's empty nested object merges nothing instead of replacing the whole section; a preset that sets nothing is refused. - A setting action with no source path fails closed. - Docs, schema, and comments say packs referenced by the global config may declare setting actions (they inherit its source path), and that confirm doesn't apply. - Move the CLI help note below the subcommand list. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Honor confirm on setting actions and explain keys containing "." A setting or settingPreset action with "confirm": true now asks before saving and shows the equivalent `cmux config` command. The project-action trust prompt never covers the global config, so this is the only prompt these actions get. A path that only resolves when one of its keys contains "." (for example a workspaceGroups.byCwd entry for ~/src/app.web) is refused with a keyContainsDot error that says so, instead of "isn't a cmux setting". Such keys stay unsupported; there is no escaping syntax. Docs, the CLI contract, and the cmux-settings skill say so. Adds a test that packs the global config references may declare setting actions while a project config's packs may not, and that confirm reaches both the palette action and the tab bar button. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Write fractional setting values in their short form JSONSerialization prints a Double with 17 significant digits, so `cmux config set terminal.scrollSpeed 1.4`, a cycle entry, or a preset leaf wrote 1.3999999999999999 into cmux.json and the CLI echoed it. CmuxSettingValue now hands JSONSerialization an NSDecimalNumber built from Swift's shortest round-trip text, and preset leaves are re-encoded through CmuxSettingValue so every path writes 1.4. CI run 36310670056 on e478696 caught this through the new commandLineDescriptions and confirmationDialogShowsTheEquivalentCommand tests; this adds a file-text assertion for set and preset. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Start setting changes from the live value; keep project configs out Review fixes: - Most settings live in UserDefaults, and cmux.json only overrides them. toggle, cycle, and get now read a key the file doesn't set from the app's UserDefaults (CmuxSettingLiveValues, backed by the setting catalog) before the schema default, so the first press after changing a setting in the Settings window flips what the user sees. The app reads its own defaults; the CLI reads the enclosing app's domain. `cmux config get` says where the value came from, and `unset` is described as removing the key from cmux.json. - A project config could override a global setting action's title, shortcut, or confirm through an actions entry or a tab bar button. Both overrides are now ignored for setting actions unless they come from the global config. - The failure alert falls back to a localized "couldn't read or save cmux.json" message for store errors that have no description. - keyContainsDot only fires when the rejoined key exists in the file or looks like a path, so a typo under a map keyed by names stays an unknown path. - `cmux config get --json` uses `path` and `file` like the other subcommands, plus `source`. - The schema no longer says confirm doesn't apply to setting actions. - The setting-action docs keys are translated into the 18 other web locales, and CHANGELOG has a line. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Only trust live values that are stored in cmux.json form Second review pass: - Some UserDefaults values aren't stored the way cmux.json spells them: app.minimalMode is a presentation-mode string, and app.keepWorkspaceOpenWhenClosingLastSurface is stored as the opposite flag. Lists and maps are often stored as text. The live resolver now maps those two explicitly and otherwise accepts only a scalar whose type the schema allows at that path (CmuxConfigSchemaPathLookup gains declaredTypes(at:)). A catalog-wide test fails if a key the resolver accepts stores a default that disagrees with the schema default, so a new transformed key has to get a mapping. - Re-setting a fraction that's already in the file is no longer reported as a change: receipts and the no-op check compare numbers in the same short form they are written in. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Map the session content width sentinel back to false The catalog test caught terminal.sessionContentMaxWidth: UserDefaults stores -1 for "no cap", which cmux.json spells false. The live resolver now maps widths under the minimum to false. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Move the changelog line to the end of Unreleased/Added Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fix setting action docs rich text Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fix setting presets from global packs Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fix settings module import Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fix config executor settings import Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Add setting actions dogfood tour Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fix setting actions tour JSON Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Refine setting actions dogfood tour Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Make setting actions tour reload config Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Isolate setting actions dogfood config Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Use isolated home in setting actions tour shell Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fix setting actions tour cleanup Use the isolated tour home when restoring config so cleanup cannot remove the user's config. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Harden setting action persistence Prefer current global presets, guard symlink retargets during publication, and keep tour cleanup isolated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ow-ai#15217) * test: predicted echo stays off for surfaces not known to be remote A local shell that echoes slower than the threshold would draw predicted glyphs, and every registered surface's output is copied and scanned while the setting is on. Both tests fail on main. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Predict echo only on remote terminals; withdraw on pasted and sent input Every terminal registered with the prediction center, and a local shell was kept out only by its measured echo latency. A shell on a loaded Mac can echo slower than the 25 ms threshold, and while the setting was on every terminal's PTY output was copied and scanned on the main actor. Each surface is now classified at its first keystroke after prediction starts for it: only a terminal whose shell runs on another machine (an SSH or Cloud machine, a remote tmux pane, another signed-in Mac) gets an active engine, and the output inbox accepts only those surfaces, so a local terminal's reads are dropped before any copy. Input that bypassed the keystroke path (a paste completing, text and keys sent over the socket, the TextBox, a paired phone) left an armed echo run in place. A key typed right after a paste was drawn where the pasted text was about to land. That input now withdraws the prediction, like an editing key does. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Turn predictive local echo on by default and move it to Terminal settings With prediction limited to terminals whose shell runs on another machine, the setting leaves Beta Features and defaults on. It is now Settings > Terminal > Predictive Local Echo and `terminal.predictiveLocalEcho` in cmux.json and the schema. It stays stored under the former beta UserDefaults key, so anyone who turned the beta off keeps it off. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Honor the catalog default for predicted echo; withdraw on remote tmux named keys The center read its setting with UserDefaults.bool(forKey:), which reads an unset key as off, so a catalog default of on would never reach it. It now takes the default from the catalog. A transport-named key in a remote tmux pane skipped the prediction path and left drawn glyphs in place. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Classify Dock terminals for predicted echo; reclassify when a remote session starts or ends A remote terminal moved into a Dock was classified local because only the workspace split tree was consulted, so it never predicted. Classification now resolves the hosting Dock first, as terminal link opening does. Classification was fixed at a surface's first keystroke, so a terminal typed into before its remote session registered (a reconnect that keeps the surface) never predicted in that runtime. A change to a workspace's remote terminal set now withdraws and classifies the surface again at its next keystroke. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: deleted text reappears after Backspace, Ctrl-U, Ctrl-W over in-flight predictions A frame-composing harness (grid plus overlay every 16.7 ms) flags any column the user deleted and saw blank that later shows text again. 6 of 8 fail. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Keep deleted predictions blank until the remote erases them A Backspace over a glyph still in flight no longer lets its echo paint the character back: the cell is drawn blank until the erase lands. Ctrl-U, Ctrl-W and Option-Backspace mask in-flight glyphs the same way, and keys whose effect is unknown (Return, arrows, pastes) raise a barrier behind the glyphs already sent instead of dropping them, so nothing vanishes and reappears. A glyph retyped onto a blank cell replaces the blank. The simulation now checks blanks at the end of the run: a blank over text already where it belongs fails only if the remote never rewrites that cell. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: Return answered by a newline suspends prediction; Ctrl-U after Return blanks the submitted line Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Resolve the barrier on the remote's newline; match Ctrl-U and Ctrl-W by character Output answering Return, an arrow or a paste (CR LF, a redraw) resolved the barrier only when printable; a newline counted as a misprediction, so fast typing followed by Return could suspend prediction. A barrier that expires unanswered guessed nothing and no longer counts either. Ctrl-U behind a barrier acts on the new line and leaves the submitted glyphs drawn. Ctrl-U and Ctrl-W are recognized by the layout's character rather than the physical key. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Trace predicted echo in dev builds while a flag file exists Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: drain-before-parse draws echoed glyphs over the cells to their left Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: zsh highlighting suspends, images keep stale glyphs, slow jittery links stop drawing, main-actor scan cost Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: links over 1.5 s never arm; one stall stops drawing until the user pauses Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: type-ahead and masked password prompts draw secrets; status repaints stop prediction Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Keep predicted echo safe at password prompts, alive through rewrites and stalls Found while dogfooding over a slow link and by adversarial tests: - zsh answers the second key of a line with BS and a rewrite of the first ("\bec"), which withdrew everything as a misprediction and left prediction dark for the rest of the line. The engine now remembers what it saw echoed on the row and accepts a move back over it followed by the same text; CSI n D reports its count for this. zsh-syntax-highlighting recolouring is the same. - A key typed ahead of a password prompt got its cooked-mode echo and armed the run, so the rest of the password was drawn. Drawing now needs two echoes in a row since the run last ended, and a `*` never counts. - An overdue echo or unmodelled cursor motion now hides what is drawn but keeps matching the keystrokes in flight, so one stall no longer leaves the line untracked while the user keeps typing. The lifetime scales to three round trips on links slower than 500 ms. - A cursor save and restore (a status line or clock repaint) is ignorable; kitty graphics and sixel images are disruptive. - Output with nothing in flight is skimmed instead of classified byte by byte, and the inbox keeps at most 1 MiB per surface, reporting a drop. - Backspace over an echo not yet painted blanks it until the erase. Blanks that output removes stay until the next presented frame, and the host drains a surface's pending output before retiring at a frame. The breaker tests that armed on one echo now arm on two. The two render race tests are disabled until ghostty tees after the parse. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: point the disabled render-race tests at the ghostty tee fix Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: read the final status outside the expectation The determinism guard reads a duration literal inside an assertion as a wall-clock check. This one is the engine's simulated clock. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: a predicted run over the layer-hosting terminal renders nothing Ghostty makes the terminal view layer-hosting. The overlay drew in draw(_:), which AppKit called with the host's whole bounds as the dirty rect, and the run never reached the screen: on a dev build with a 1.8 s link, ten predicted glyphs sat drawn, unhidden and on top for over a second with nothing visible. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Render the predicted run into the overlay layer's contents The terminal view is layer-hosting (ghostty's IOSurface layer is its layer), and AppKit's draw pass for a subview of it never put the run on screen. Render the run into a bitmap of the overlay's own size and set it as the layer's contents from updateLayer, which does not depend on that pass. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: a layout hold with no presented frame freezes the overlay On a dev build the rendered-frame callback never fired: it is counted in the Metal layer's nextDrawable, and ghostty presents through its own IOSurface layer. The hold set when an erase removed blanks was never released, so the overlay kept a blank over a cell the grid had already redrawn, and hid every later prediction on that surface. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Release the layout hold when no frame arrives The hold waits for a presented frame, and the host may never report one. Release it after confirmationHold, the same bound a held confirmation has, and include that deadline in nextExpiry so the host's timer re-anchors the overlay. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Add a predicted-echo recording script for dogfood GIFs Puts a local sshd behind a delay proxy, opens a cmux ssh workspace on a tagged build, types paced keys through the debug socket and records the window as a GIF. The dogfood tour format cannot express a slow link, a remote workspace, timed keys or video. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix predictive echo for consumed bindings * pin Ghostty tee fix release * fix predictive echo binding classification * classify Ghostty local-only bindings * fix Ghostty binding flag typing * record local-only GhosttyKit checksum * rebase Ghostty binding flag onto main * pin Ghostty local-only binding fix * record rebased GhosttyKit checksum * rebase Ghostty binding fix onto current main * fix Ghostty gitlink * record current GhosttyKit checksum * retain GhosttyKit checksum after main merge * ci: record rebased GhosttyKit checksum --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ai#16140) * test(ci): update-homebrew must keep triggering-run metadata out of scripts update-homebrew.yml substitutes github.event.workflow_run.head_branch and the dispatch input into its `run:` script as text. The test renders the version step the way the runner does, with a hostile branch name, and checks nothing executes; it also requires the gate to accept only the real release workflow run from a tag push in this repository. Found by a Codex Security audit (both models). Red: python3 tests/test_ci_homebrew_untrusted_input.py Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(ci): update-homebrew takes run metadata through env and trusts only the real release run Makes tests/test_ci_homebrew_untrusted_input.py green (previous commit). Root cause: the version step substituted github.event.workflow_run.head_branch and the dispatch input into its script text, and the gate accepted any run of a workflow with the release workflow's display name. The values now reach the script as env vars (INPUT_VERSION, HEAD_BRANCH, EVENT_NAME), and the gate requires the run to be .github/workflows/release.yml, from this repository, for a push. Dispatch keeps working as before. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
manaflow-ai#16148) * ci: survive an owned Mac whose Homebrew prefix the runner does not own Run 36752535712 built the app-host and UI test product on cmux-austin-mini-1-glaeda-1, 17 minutes, then lost the whole job in the "Install tmux" step: Error: /opt/homebrew/Cellar is not writable. You should change the ownership and permissions of /opt/homebrew/Cellar back to your user account: sudo chown -R cmux /opt/homebrew/Cellar brew refuses outright when the prefix belongs to another account, so the step exits 1, "Resolve selectors against the built tests" is skipped, and the job reports TEST_RESULT=failed with TEST_SUMMARY and TEST_OUTPUT both empty. The dispatch looks like a test failure and names no cause short of reading the raw log. scripts/ci/brew-ensure.sh retries the install as the prefix owner, and when even that cannot produce the command it fails with the runner name and the package to provision. Both brew installs in the E2E action use it. Each call site keeps its old inline path behind an -x check, because this action runs against the tested revision's checkout, which may predate the script. This does not provision that mini; it stops one missing package from throwing away a build that already succeeded, and makes the next occurrence legible from the step summary. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * ci: reach the helper from the workflow revision and classify its failure Review of the first commit found five ways the fix would not have helped the run it was written for. The helper was called workspace-relative, so it came from the tested revision. test-e2e.yml checks the action out of the workflow's own revision into .e2e-workflow with a sparse-checkout that did not list the helper, so a re-dispatch of the ref from the failing run, and every step of a regression bisect into older history, would have taken the old inline brew install and failed identically. Add the helper to both sparse-checkout blocks and prefer the .e2e-workflow copy at both call sites, the way the frames script already does. Guard on -f and invoke through bash instead of -x. A lost mode bit made the ffmpeg branch fall through to an unconditional brew install, which fails on an unowned prefix even when ffmpeg is already there: worse than no helper at all. The retry now runs sudo -n, gated on sudo -n true, which is how the rest of this action and run-in-console-session.sh treat passwordless sudo on these Macs. Without -n, a host with a controlling tty would block on the prompt until the job timed out, holding a Mac. HOMEBREW_NO_AUTO_UPDATE was a command prefix on the first install only, and sudo's env_reset would drop it regardless, so the retry could trigger a full brew update inside a step that had already burned the build. Carry it over the sudo hop with /usr/bin/env, as action.yml:545 does. Every failure path now prints [cmux-ci machine: brew-provision], so machine_failure.py classifies it directly instead of depending on Homebrew's own wording reaching the log. A root-owned prefix gets its own message, since brew refuses to run as root and no hop fixes that. Verified: shellcheck and bash -n clean, actionlint clean on test-e2e.yml, both YAML files parse, tests/test_ci_machine_failure.py passes with the new case (8 tests), tests/test_ci_self_hosted_guard.sh passes, and all eight branches of the script exercised under brew/stat/sudo shims exit 0 only with the command present and non-zero only with it missing. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(ci): pin the new helper in the e2e sparse-checkout guard test_the_build_runner_runs_the_tests_so_a_run_queues_once asserts the sparse-checkout list verbatim, so adding scripts/ci/brew-ensure.sh to it reddened app-host-execution. The addition is intended: without it the helper is absent from .e2e-workflow and the action falls back to the tested revision's copy. Also assert every entry exists. A sparse-checkout of a path that is not in the repo is silent, so a typo there would leave the action reading a missing file and silently taking the older copy instead. Verified: python3 tests/test_ci_e2e_compilation_cache.py, 25 tests OK. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…#15887) * test: cover agent activity after turn completion * fix: keep agent status running while hooks show activity * fix: ignore ambiguous late tool activity * Add subagent and waiting agent states to the sidebar The compact status glyph had one "running" state, so a pane running a fan-out of subagents, a pane parked on a background command, and a pane typing a reply all looked identical. Two of those are worth telling apart: subagent work is the loudest thing an agent does, and a pane waiting on a deterministic wakeup is not asking for anything. Claude's hooks now report what a running pane is running on through a new `set_status --work=running|subagents|waiting` option: - PreToolUse with `tool_name` of `Task` reports subagents. A Task call blocks the parent inside the tool until its subagents finish, so the state holds for exactly that span and the next parent hook clears it. No counter to drift. - Stop with a live background task or scheduled wakeup reports waiting instead of running. A re-entrant Stop stays running: that is the agent itself still going. The work state rides alongside the agent lifecycle rather than inside it. A waiting pane keeps reporting a running lifecycle on purpose, so hibernation can never SIGTERM live background work; the work state is presentational only, and the resolver reads it before the lifecycle branch. Waiting wins only when every agent in the workspace reports it, so one agent still working keeps the row running. Glyphs: subagents is a pulsing gray connected-points symbol, waiting is a still gray hourglass. Waiting does not pulse, because the agent is parked and a pulsing hourglass would claim otherwise. Both are configurable through `sidebar.compactStatusIcons`, and both reach the non-compact rows through the icon the hook sends. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> (cherry picked from commit b4bee23) * Add dogfood tours for the new sidebar agent work states Two tours over the same five workspaces (subagents, waiting, running, needs input, idle): one with the compact glyph on, one with it off so the metadata rows show the icons the hooks send. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> (cherry picked from commit 5904e8a) * Fix the work-state compile breaks and the shared-key waiting regression CI on 5904e8a caught two compile breaks the branch shipped with: the two `shouldReplaceStatusEntry` call sites in SidebarOrderingTests never gained the new `workState` argument, and a new control-socket test called `hasPrefix` on an optional response. `everyIconSlotHasADistinctState` also still pinned 11 icon slots against the 13 the branch now has. The cmuxTests target could not build, so none of the branch's own tests ran. The review that ran alongside it found three behavioral defects: Subagents never appeared on a current Claude Code. The PreToolUse row matched only `tool_name == "Task"`, and 2.x sends `Agent` for the same spawn. Both names now count, the way `AgentChatSessionRegistry.isTaskSpawn` already handles it for the mobile child-run tracker. An hourglass could cover a pane that was still working. Status entries are keyed per workspace while lifecycle states are keyed per panel, so two Claude panes in one workspace share one `claude_code` entry and the second to report wins. Waiting now also requires that every running lifecycle is covered by a waiting report, so a sibling pane mid-tool-call keeps the row running. Two panes both waiting under one key read as running, which is the conservative direction. The work state is now listed by `list_status` and `sidebar_state` as `work=<state>`, so the state behind the glyph is observable instead of screenshot-only. Also: the doc comment promised that an unknown work state degrades to a plain running row, while the socket rejects the whole `set_status` the way it already rejects an unknown `--format`; the comment now describes what the code does. `SidebarAgentWorkState.parse` dropped a `_`/`-` pass that no input could reach and a singular `subagent` alias the socket rejects, so the two parses accept the same set. The glyph header and docs/configuration.md listed Running above Waiting while the resolver checks Waiting first. Tests: the renamed spawn tool, the two-pane shared-key case both ways, two agents both parked, the listing line, and a pin on the raw values the sidebar and control-socket copies of the wire contract share. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> (cherry picked from commit 1bad27d) * Append the work option last so it does not split an older command prefix CI caught this on the app-host lane: CLINotifyProcessIntegrationRegressionTests.testClaudePromptSubmitFrom NewSessionCanReplaceStoppedSession asserts the prompt-submit command as a prefix through `--tab=`, and `--work=running` was being inserted between `--color=` and `--tab=`, so the prefix no longer matched. Three assertions in tests/test_claude_hook_clear_running_status.py use the same contiguous fragment and would have failed on their own lane for the same reason. None of those four assertions is about work states; they check that prompt-submit sets Claude running on the right tab. Options are order-independent on the wire, since the coordinator reads a parsed option dictionary, so the new optional one goes at the end of the command instead and the older assertions stay intact. Updating them to expect `--work=running` would have coupled four unrelated checks to this feature and broken them again the next time the work state for prompt-submit changed. Pinned by a new test in ClaudeHookWorkStateTests: the running command must still start with the historical prefix and must end with the work option. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sidebar): list the work option in the socket help output The `help` text for `set_status` was the one place that still omitted `--work`, while the usage and error strings in both coordinator copies already list it. Pin the work-state ordering test through the workspace id, so it stands in byte for byte for the prefix the older suites assert, and say in the comment why order independence holds: every option here is `--key=value`, which a future bare flag would not be. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Keep the sidebar pill on Running for a re-entrant Claude Stop Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: replay deterministic wait tools in sidebar hooks * fix: show deterministic Claude waits in sidebar * fix: reopen identityless fresh tool activity * fix: gate identityless activity by hook phase * test: model fresh identityless tool resolution --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…#15409) * Add the window recording request and frame model to CmuxFoundation A clip needs its parameters settled before any capture starts: the format and its default frame rate, the scale and width caps, the crop rectangle in window points, the file name, and how the frame size follows a window that is resized mid-clip. None of that needs a window or a screen, so it lives in CmuxFoundation with tests instead of inside the capture session. WindowRecordingFrameGeometry aspect-fits each captured image into the frame size chosen at the start, so a resize letterboxes rather than stretches, and a crop is adopted once and then kept for the rest of the clip. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Add cmux record for capturing a cmux window to mp4 or gif An agent working in its own cmux can describe what it changed, but it cannot show it. `cmux record start` films a cmux window, or a region of one, and writes an mp4 or a gif that can go straight into a pull request. `cmux record note` drops a caption into the clip as it runs, so a reader can follow what was being done. Capture uses ScreenCaptureKit's own-process content, so only cmux's own windows are reachable and no Screen Recording permission is asked for. Encoding is AVAssetWriter for mp4 and ImageIO for gif; frame times come from when each frame was sampled, so playback matches what happened. A recording stops itself at --max-seconds, and one runs at a time, so an agent that goes away cannot leave a capture running. The window.record.* socket methods run on the socket worker and are not main-thread callable: the window being filmed has to keep drawing while the sampler runs. They are deliberately absent from the remote relay allowlist, which stays default-deny, because a clip is local screen content rather than an object of the peer's workspace. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Add failing tests for the window recorder's crash and early-stop paths A review of the recorder found seven defects the existing tests missed. These are the reds, before the fixes: - `--fps nan`, `--fps inf` and `--fps 1e30` reach `Int(Double)` in the request decoder and trap, which takes the whole app down with the socket. The package suite now crashes with "Double value cannot be converted to Int" instead of finishing. - a gif whose recording stopped long before its frame budget cannot be finalized, because the budget is handed to ImageIO as the number of images the file will contain, so the documented `--gif` flow leaves no file at all. - `record stop` after a clip reached its own `--max-seconds` limit reported "no recording is running" rather than the finished clip. - an unknown recording id, and a note for a clip that already stopped, had no coverage at all. - the mp4 writer's no-frames path was not checked for leaving a stub file behind, and the two-frames-one-tick test only asserted a nonzero duration rather than the timestamps it exists to pin down. - `fit` upscales a window that shrank mid-clip, which its own name says it does not do. `remember` on the registry stops being private so a test can set up a finished clip without a window on screen. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fix the window recorder's crash, race and early-stop paths Seven defects a review found, in the order they bite a caller: - `--fps nan` and `--fps 1e30` trapped in `Int(Double)` and took the app down. The decoder now refuses a value that is not finite or does not fit an `Int`. - a gif was created with the recording's frame budget as its declared image count, so a clip stopped before that budget could not be finalized and left no file at all. The count is a floor rather than a cap, and how many frames a clip ends with is only known when it stops, so declare one and add as many as were captured. - `stop` could close the file while an append was suspended inside the session actor, which is an uncatchable AVFoundation exception for the mp4 writer and concurrent state for the gif writer. Appends and the close now take turns. - two `record start` calls could both pass the one-at-a-time check, because the check was followed by the await that opens the capture. The slot is claimed before that await. - `record stop` after a clip reached its own `--max-seconds` limit said no recording was running; it reports the finished clip, as `status` already did. - the mp4 writer cancelled a writer it had never started when a clip captured no frames, which raises rather than returns an error. - the output file was deleted before the first frame was written, so a failed recording destroyed whatever was at `--out`. Frames go to a hidden sibling file and only move into place once the clip closes, a directory or device at that path is refused, and a partial file is cleaned up on every failure path. Also: a window closed mid-clip reported an empty rectangle and was captured as a one-pixel frame for the rest of the clip, so an empty rectangle ends the recording and keeps what came before; `fit` no longer magnifies a window that shrank mid-clip, which is what its name always claimed; and `window.record.*` is listed in `system.capabilities` for local callers, where the relay policy still filters it out. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Refuse a record flag that swallowed the next flag as its value `cmux record start --label --gif` named the clip "--gif" and recorded an mp4, because the shared option parser takes whatever follows a flag. The record command now hands a value that looks like a flag back to the unexpected-arguments check, which names both, and the positional recording id for `stop`/`status` no longer accepts one either. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Map an unusable recording output path to invalid_params `outputNotAFile` was added to `WindowRecordingSessionError` in the fix commit for the output-path handling, but the router's switch over that enum was never extended, so the app target stopped compiling: Sources/TerminalController+WindowRecording.swift:175:13: error: switch must be exhaustive Nothing on the pull request caught it. The checks that run on a pull request build the packages and the tooling, not the app target, so this first appeared in a dispatched UI test run: https://github.com/manaflow-ai/cmux/actions/runs/36412249068 The missing case is `invalid_params`: `--out` naming a directory, a device or anything else the recorder may not replace is the caller's parameter, not a cmux failure, and an agent that reads `internal_error` retries instead of fixing its flag. The mapping had no test at all, which is why a missing case could sit here. `recordingErrorCode(for:)` is now internal and every case of the three error types it understands is pinned, including the unrecognized fallback. A compile error cannot be committed as a red test first, since the test would not build either; the build log above is the red evidence, and the test is what keeps the codes from drifting. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Share the recorder's parameter decoding with a screenshot request `window.screenshot` takes the same region, the same scale and width caps and the same rule about an output path as `window.record.start`, and a caller who finds a rectangle in a clip should be able to shoot it with the same four numbers. Written twice, the two would drift: one would round a region's edges differently, or accept a relative path, or trap on `--max-width 1e30` where the other refuses it. So the decoding moves into `WindowCaptureValueDecoding`, which both requests translate into their own public failures. The recorder's API and every message it produces are unchanged; `WindowRecordingRequest.make` now decodes the format first and maps the shared failure back into its own, which is why the format still names itself in a path error. `WindowRecordingFrameGeometry.plan` gains an overload taking a region, a scale and a width quantum rather than a recording request, so a still can plan its crop with the code a clip uses. H.264's even-width rounding becomes that quantum instead of being read off the format. `WindowRecordingLabel` takes a fallback, so a label that sanitizes away to nothing is filed as a screenshot rather than under the word "recording", and `WindowRecordingOutputNaming` gains the screenshot directory the DEBUG command has always used plus a filename overload taking an extension. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Discard a timed-out record start and keep standalone -- in notes A record start that timed out on the socket kept running and could register a recording the caller had been told failed. The registry now tracks each start by token; on timeout the socket handler abandons it and waits for the session to release its writer and partial file before answering. A start that is stopped or abandoned mid-capture no longer opens a writer, and fail() releases a writer even after the session reached a terminal state. record note now strips only a leading -- terminator, and cli.usage.record carries every catalog locale. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Scope recording limits and output naming onto WindowRecordingRequest The package conventions lint rejects all-static public types, so the limits and the default output naming become extensions on the request. The mp4 test names its encoded clip length plainly so the determinism check does not read it as a measured wall-clock time. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Add `cmux shot` for screenshotting a cmux window An agent that films a window with `cmux record` still has no way to grab one frame of it. A clip is the wrong artifact for "what does this sheet look like now": it is slower to produce, heavier to attach, and a reviewer cannot read a label off it. `window.screenshot` captures one of cmux's own windows to a png or a jpeg, through the same ScreenCaptureKit path the recorder uses. No Screen Recording permission is involved, so it works on a release build and inside CI, where the DEBUG-only `screenshot` command does not exist at all. The two commands share every decision they could disagree about: parameter decoding, crop and scale planning, frame capture, output-file handling and window resolution. `cmux record --region 0,0,420,900` and `cmux shot --region 0,0,420,900` frame the same rectangle of the same window, so a detail spotted in a clip can be shot with the same four numbers. The image is encoded beside the output path and moved into place, so an existing file there is replaced only once there is a complete image to replace it with, and `--out` pointed at a directory comes back as `invalid_params` rather than deleting it. Local socket only: `window.screenshot` is not on the `cmux ssh` relay allowlist, since it reads pixels of windows the remote session does not own. ## Changelog Added: `cmux shot` captures a cmux window, or a region of one, to a png or a jpeg. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Compare recorded pixels through an Equatable value Optional tuples are not Equatable, so the caption test did not compile. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Add a cmux-capture skill for screenshots and clips `cmux shot` and `cmux record` exist, but an agent looking for a way to show a reviewer what a change looks like has to find them first. The skill says which of the two fits a still and which fits a change over time, how to reuse one `--region` rectangle across both, and what an agent recording its own cmux needs to know: the recording stops itself, one runs at a time, captions are drawn in as it goes, and nothing but cmux is in the frames. `references/commands.md` carries what `--help` cannot: the response fields, the error codes and the limits the app enforces. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Add a capture topic to cmux docs `cmux docs` is where an agent with no repository checkout looks for a topic, so the capture skill and its command reference are reachable there too, and from the agents topic next to the hook integrations. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Guard the capture skill against drift from the code A doc that names a response field or a state the app no longer emits sends an agent after something that is not there. Two such errors were in the draft: it called the recording states `recording` or `stopped` when they are `recording`, `finished` and `failed`, and called the screenshot `format` field `png` or `jpg` when the response carries `jpeg` and only the file is named `.jpg`. Both came from reading the CLI and not the payload. The guard reads both sides: documented flags against the CLI help, documented states and formats against their enums, documented response fields against the keys the payload actually sets, and documented error codes against the ones the socket methods return. It needs no build, and it runs in the skill contract workflow when the skill or the code behind it changes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Probe the capture docs topic from the CLI contract `test_cli_contract_help.py` runs every probe line against the built CLI, so the new topic is checked the way the dock topic is. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(skills): complete capture discovery contract * fix(skills): advertise capture docs in CLI help * Fix window capture edge cases Validate and safely clamp region geometry, including origin-anchored crops and finite values beyond integer range. Build GIFs only after their exact frame count is known, staging compressed frames on disk. Keep screenshot output promotion on the waiting socket worker so a timed-out capture cannot later overwrite the requested path. * fix(capture): preserve directories at promotion * fix(capture): bound frame work and caption history * fix(capture): isolate uncooperative frame requests * fix(capture): bound GIF staging and sampling * Pin the capture doc claims a reviewer found wrong Nine doc findings from a review of this skill, and three guards so the same class of drift is caught next time rather than read past. The corrections: the error table said `invalid_params` covered an unknown record subcommand, which never reaches the socket; a region under 8 or over 100000 points, a relative `--out` and an `--out` whose extension contradicts the format were all undocumented; `not_found` covers an unknown recording and a window that was never open, not only one that closed mid-capture; `conflict` also answers a `stop` that raced its start; `fps_effective` needs two frames spanning a nonzero time, not just two frames. The guards: the documented region bounds are compared against the `Double` constants both requests enforce, every relative link in the two pages has to resolve to a file and to a heading that exists, and every Swift file these checks read has to be in the workflow's trigger paths. The last one was missing `WindowRecordingRequest.swift`, so a change to the recorder's own bounds could not have failed this test. The link guard immediately caught a forward reference: the skill pointed at `dogfood-scenarios.md#record-a-clip`, a heading that arrives with the dogfood record step, so the anchor is dropped until that lands. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Wait on deadlines and make the sample schedule a value Two CI guards refused the capture branch, and both were pointing at something worth changing rather than at the gate being strict. `WindowScreenshotWriteTests` waited by yielding a fixed number of times. N yields is however long N reschedules take, so the wait tightens exactly when the runner is busy, which is when a passing test turns into a flake. It now yields until a five second deadline instead. `WindowRecordingSampleSchedule` was an enum of static functions that took the previous target back in as a parameter, so the caller held the state and the type held none. It is now a struct the sampler owns and steps forward, which is also why its tests can assert on `targetUptime` after a served slot rather than only on a returned number. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Pin the writer's error codes, which nothing tests any more `record start` opens the output file before it answers, so an `--out` the recorder cannot create fails inside the writer rather than in the request decoder. `TerminalController.recordingErrorCode(for:)` has no branch for `WindowRecordingWriterError`, so every one of them falls through to `internal_error`, and an agent told by the docs to fix its flags on `invalid_params` reads its own unwritable path as a cmux bug and retries. This test covered it until the branch was repaired; the case that made it redundant, `stoppedWhileStarting`, is gone, but the writer errors are not. It is restored first, failing, and the mapping follows. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Tell a caller when the recorder could not open their clip `WindowRecordingWriterError.setup` is the file the caller named refusing to be created, so it answers `invalid_params`; a frame or a close that fails is the encoder, so those stay `internal_error`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Pin the limits and codes an agent reads, before fixing them Three claims the capture docs make that nothing checked. `invalid_params` tells an agent to edit its flags, so it cannot also mean "this window has nothing in it": no flag makes an unrendered window capturable, and the mapper sent all three geometry failures down that branch. The test now pins the split. The sample schedule counted one slot past every elapsed slot, so a sampler arriving exactly on a later slot threw that frame away and waited a whole interval more for the next one. The new case asks for the slot it is already in time for. The limits table was checked for flag names and never for numbers, which is how it came to promise a gif `--max-width` of 4096 against a ceiling of 1280, and to omit the two gif budgets that have no flag of their own. The guard now reads the constants and asserts the page states each one, and the region check compares values instead of mirroring a source line. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Keep an on-time frame, and stop blaming flags for an empty window `recordingErrorCode` now answers `not_found` for `emptyWindow`, beside the two geometry failures a parameter can fix. An agent that reads `invalid_params` changes its flags and tries again, which no flag would have helped here. The schedule rounds up to the first slot at or after now instead of adding one to the elapsed slots, so a target that is exactly `now` is served rather than skipped, and it steps once more in the one case where the division rounds down in binary. Two comments claimed a window resized mid-clip never ends a recording. That holds for the frame size, which is fixed when the clip opens, but not for a `--region` the window has shrunk out of: the crop is planned again per frame, planning throws, and the clip ends in state `failed` with the frames it had. Both comments now say so. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Document the gif ceilings, and offer an index in --help The limits table promised a gif `--max-width` of 4096 while a gif stops at 1280, and said nothing about the two budgets a gif has and an mp4 does not: 960 frames, and 4000000 pixels a frame. The reference page owns the limits and `--help` owns the flag list, so the ceilings go here rather than into a wider flag column. Also corrected on that page: a region a shrinking window leaves behind ends a clip, which the region section had not admitted; `conflict` no longer covers a stop that raced its start, a case that no longer exists; an `--out` under a read-only volume is `invalid_params`; a window with no capturable content is `not_found`; and `rpc window.screenshot` takes a window id, because the CLI is what resolves a ref or an index. The skill page said `record status` reports the effective frame rate, which only `--json` carries. `--help` for both commands listed `--window <id|ref>` while the usage line and the reference page both offer an index, and the CLI has always accepted one. Same token in all nine locales, since it is a spec rather than prose. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(recording): stop asserting how many samples the encoder writes `twoFramesInTheSameTickStillGetIncreasingTimes` expected a two frame clip to read back as exactly two samples, and the CI VMs write six: with no hardware scaler available, the encoder decides how many samples a clip of two frames a 1/600 second apart ends up with. The property under test survives. A dropped frame, which is what an equal presentation time used to cause, leaves a single sample, so a floor of two still catches it, and the times still have to start at zero, increase, and put the second sample a tick rather than a frame interval after the first. The expectations now print the times they read, so the next surprise from the encoder arrives with its evidence. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(recording): pin that a gif keeps the frames staged before its limit A gif that reaches maximumStagedBytes mid-clip currently loses every frame: finish() retries the frame it could not stage, throws again, and never creates the destination, so the session deletes the working file. That contradicts fail(_:), which promises "Keep whatever was captured before the failure". Red before the fix: the finalized gif does not exist. ## Changelog none Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(recording): pin what a mid-clip crop replan may and may not fail on A clip's encoded frame size is fixed by its first frame, so replanning a later frame with the gif pixel ceiling still applied ends a recording over a size nothing encodes: a window that grows mid-clip fails with gifFrameTooLarge even though the frame is drawn into the opening size. A region the window shrank out of is the one resize that still has to end the clip. Red before the fix: planCrop does not exist. ## Changelog none Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(recording): keep a gif's staged frames when the last one cannot stage `finish()` staged the frame it was holding without clearing `pending`, so a clip that had already reached its staging limit threw again from the one call that is supposed to finalize it. The session's `fail(_:)` promises the caller "whatever was captured before the failure", and that promise was being broken for exactly the recordings that need it. Drop the one frame that has no room and finalize the frames before it. A clip with nothing staged still throws, so a recording with no frames still reports why it has none. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(recording): re-crop a window that grows mid-clip instead of ending the gif The per-frame replan called the first-frame planner, which applies the gif pixel ceiling. A window that grew past that ceiling mid-clip therefore ended the recording over a frame size nothing encodes: the encoded size was fixed by the first frame, and the grown crop was only ever going to be drawn into it. `planCrop` plans the crop with no pixel ceiling, and the session uses it for later frames. The one resize that still ends a clip is a `region` the window has shrunk out of, because no rectangle is left to sample. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(cli): give the `screenshot` alias its contract row `check-cli-contract-verbs` reads dispatched verb strings, and the shot arm dispatches both `shot` and `screenshot`. Only `shot` had a row, so the guard failed the whole run with "top-level verb screenshot is in no command table". The first cell may name a verb and its aliases, so both go in the one row. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: validate screenshot requests and capture docs --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: reproduce answered agent notification ring Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: clear answered agent notification rings Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: clear Codex answered notification rings Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: fence Codex prompt notification retirement Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: normalize agent session fallback matching Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: stamp ordered Codex progress hooks Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: limit Codex progress stamps to safe hooks Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: clear agent rings on terminal input Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: clear agent rings in dock terminals Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: remove unused semantic notification binding Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: isolate agent notification mutation fixtures Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: clear queued mutations between notification fixtures Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: fence superseded agent prompt retirement Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: preserve newer same-session prompts Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: index unread agent prompts Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…anaflow-ai#15989) * test: reproduce late Agent Chat startup stealing session focus * fix: keep late Agent Chat startup replies from changing selection * Make Agent Chat fork and handoff actions recoverable (manaflow-ai#15994) * test: reproduce stuck Agent Chat fork and handoff lifecycle * fix: make Agent Chat fork and handoff actions recoverable * fix(agent-chat): deduplicate reconnecting session actions Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(agent-chat): keep main's pinned tsc in the check script The main merge resolved agent-chat/package.json toward this branch's `bun x tsc --noEmit`, but main had since pinned typescript 5.9.2 as a devDependency and added tests/test_ci_guard_workflow_structure.py, which asserts the exact check script. CI fast guards failed on that assertion. The pinned dependency came through the merge, so ./node_modules/.bin/tsc exists; only the script string was stale. Verified: python3 tests/test_ci_guard_workflow_structure.py passes, and `bun run check` in agent-chat is 106 pass / 0 fail across 18 files. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…w-ai#14852) * Checkpoint terminal scrollback so a crash no longer loses it The 8 s session autosave never captures scrollback, so after a crash or SIGKILL every terminal restored empty; only clean quit, power-off and update relaunch persisted it (manaflow-ai#2016, manaflow-ai#2194). Add slow, bounded scrollback checkpoints next to the primary snapshot: - at most every 60 s, only after 5 s without typing, driven from the existing autosave timer; - only terminals that produced PTY output since their last capture, flagged by one relaxed atomic load per PTY read in the existing tee; - at most 3 Ghostty VT exports per checkpoint, one per main-queue turn, stopping early on typing or after 50 ms of main-thread capture time; - truncation, encoding and writes on a utility queue, one file per terminal under session-<bundle>-scrollback/, pruned to live panels; - the same eligibility gates as the quit path (running command, hibernated agent), with stale checkpoints deleted. After an unclean exit, startup restore fills each terminal's missing scrollback from its checkpoint; the newer of snapshot and checkpoint wins, and the restore path still applies its own replay gates. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Bound checkpoint capture cost and keep restored scrollback across crashes Review follow-up for the scrollback checkpoints: - Split the capture: only Ghostty's VT export runs on main; reading the export file, CRLF normalization, the 4000-line tail, truncation, encoding and the write run on the utility queue. A terminal whose export alone exceeded the 50 ms budget, or failed, is skipped for 10 minutes. - Seed checkpoints from restored scrollback when a restore completes, keyed by the restored panel ids and without a VT export, so a second crash before the next checkpoint keeps it. - Record scrollbackCapturedAt on snapshots saved with scrollback. A newer scrollback-bearing save wins over a checkpoint even when it deliberately omitted a terminal's scrollback; the 8 s autosave (no marker) still yields to checkpoints. - Re-mark a terminal pending when its export read or file write fails. - Keep a terminal pending for one more checkpoint when output arrived between planning and capture, since the PTY tee runs before Ghostty parses those bytes. - Disable checkpoints under automated test runs, like session restore. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Discard checkpoint exports on quit and seed only after a crash - A checkpoint interrupted by quit or restore now deletes the export files it already wrote synchronously on main (one unlink each) and leaves those terminals pending, instead of handing them to the utility queue, which may not run before exit. - Seed checkpoints from restored scrollback only after an unclean previous launch, not on clean launches or manual reopen, to avoid rewriting up to 400 KB per terminal needlessly. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fix overlapping access in the scrollback checkpoint merge Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Discard checkpoints after a clean exit so a later crash cannot restore them A clean quit saves scrollback into the snapshot but left the checkpoint files behind until the next launch's first checkpoint pruned or rewrote them. Restore reuses snapshot panel ids (workspace and dock terminals), so if that next launch crashed first, its 8 s autosave (no scrollback, no capture marker) matched the old records and the crash restore replayed the earlier launch's scrollback over what the quit saved. Startup now deletes the checkpoint directory when the previous launch exited cleanly, before the restore decision so a skipped restore also drops them. After an unclean exit the checkpoints are that launch's own and are still merged and kept until the restored terminals are seeded. Moves the checkpoint coordinator setup into the AppDelegate checkpoint extension to keep AppDelegate.swift's growth down. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Inject one scrollback checkpoint activity owner and delete removals before returning The PTY tee bridge, the checkpoint coordinator and the persist step now share one `TerminalScrollbackCheckpointActivity` owned by the composition root (`GhosttyApp.terminalScrollbackCheckpointActivity`) instead of a `.shared` singleton, so a caller cannot hand the coordinator a different instance than the one the tee writes. Planned removals (a panel that stopped being eligible) are now applied with `queue.sync` after earlier queued writes, and only captures and pruning go async. A crash before the async block ran used to leave the old checkpoint on disk for the next unclean restore to merge. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Test scrollback checkpoint output interleaving Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Make scrollback checkpoints generation-safe Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Preserve scrollback for idle agent sessions Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Match checkpoint agent liveness policy Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ai#15928) * test: cover Ctrl-letter key event encoding Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: encode synthetic Ctrl-letter keys Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: compare named Ctrl-letter bytes Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: report named key releases Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: avoid shadowing Ghostty key event type Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: retain synthetic key press capture Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: match Ghostty legacy Ctrl-I and Ctrl-M Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: restore bonsplit pin for macOS compile Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
manaflow-ai#15420) * Move browser and tmux OSC filter tests into their package test targets BrowserSystemProxyMirror, BrowserUserAgentPolicy, BrowserNavigationDecisionHandler, BrowserHiddenWebViewDiscardManager and BrowserHTTPBasicAuthPromptCoordinator live in CmuxBrowser, and RemoteTmuxNotificationOSCFilter in CmuxRemoteSession (manaflow-ai#13787), but their 68 tests compiled into cmuxTests and ran inside the launched app host. They use only package API, so they move unchanged except for dropping the cmux import. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Drop the blank line the removed import block left Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Keep app shard metadata aligned with package tests * test: give split fixtures deterministic geometry --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add failing tests for agent wake verification The tests need new types, so this commit also adds minimal stubs: AgentWakeVerificationState never changes state, and the Workspace entry points (beginAgentWakeVerification, failAgentWakeVerification, dismissAgentWakeFailure) do nothing. The state machine and Workspace tests fail against these stubs. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Check that a woken agent came back Waking a hibernated agent types its resume command into a fresh shell, but nothing checked whether the agent started. A missing launcher or a gone session left the pane at a shell prompt with no sign of failure. A wake now starts a check for the pane. An agent hook reporting a PID or a lifecycle state counts as success. The check fails when the resume command returns to the prompt first, or when nothing reports within 90 seconds and no live agent process is found. A failure shows a banner on the pane (Retry, Show command with Copy, Close), an "Agent didn't resume" sidebar row, and one notification feed entry. The row is kept out of the saved session. Closing the pane, hibernating it again, or a later agent report clears the check. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Only fail quick resume exits and never retype into a running program A resume command that returns after 20 seconds ran the agent, which the user then quit; for agents without hooks that was reported as a failed wake. It now counts as resumed. Retry is hidden, and refused, while a command still runs in the pane, since the text would go to that program. The failed-wake row survives a sidebar reset, and a retired workspace skips the deadline. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Confirm wakes only from the woken agent's own evidence Hook reports now count only when they come from the woken agent and name their pane explicitly; the focused-pane fallback no longer confirms a wake. A resume command that ran a while is no longer taken as proof: a pending check samples the pane for a live agent process every few seconds instead, and a command that ends before any confirmation fails the wake. The verification deadline constants are nonisolated, which fixes the new Swift warning. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(cloud): cover the reads that find a Cloud machine gone A status read and an access preflight both move the row to `destroyed` themselves when the provider no longer has the machine. Once either does, `destroyVm` can never see the row again (its lookup skips destroyed rows) and the provider-status cron skips it too, so the model-plane revoke and the `vm.destroyed` usage event that the cron performs for the same transition have to happen at these write sites as well. Fails today: both retire the row and record nothing. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(cloud): finish the destroy work where a machine is first found gone Three places retire a Cloud machine's row after the provider says it no longer has the machine: the status read, an access operation's resume preflight, and the stats read. Each wrote `status = destroyed` and stopped there, while the reconcile cron doing the same transition also revoked the machine's model-plane tokens and recorded a `vm.destroyed` ledger event. That difference was permanent, not a race to lose. `destroyed` is terminal: `findUserVm` hides such a row from every destroy request and `reconciliationCandidates` drops it from the cron, so whichever of the three got there first left a machine that never appears as destroyed in the ledger and never has its route tokens marked revoked, with nothing able to finish the job afterwards. All three now go through one `applyObservedProviderStatus` that performs the write and, when the write lands on `destroyed`, the same revoke and ledger event as the cron. The status route hands `getVm` the model-plane revoker it already builds for delete. The new `provider_status_*` destroy reasons join the analytics allowlist so the event says which read noticed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(cloud): derive the observed status inside the shared destroy path Review found that the previous commit took the row's new status from each caller. That let the access preflight and the stats read hardcode "destroyed" for a provider 404, including for a machine with a persistent home volume, which observedDbStatus maps to "paused" because the compute is gone and the machine is not. Those rows were terminalized and billed a vm.destroyed that never happened, and nothing revisits a terminal row to take either back. The status is now derived inside applyObservedProviderStatus, so every entrypoint agrees about what a 404 means. reopenBaseIfProviderDeleted is the one caller that must override it, and passes forceStatus with the reason: its row is a Base's active generation, and leaving it paused would hand the same dead provider id back on every later open. Also threads the model-plane revoker through getVmStats from its route, covers the stats entrypoint including the home-volume case, and fixes the two test-file regressions the previous commit shipped. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * vm: say what a paused 404 row means for credentials The access preflight carried a comment claiming a revoke was pointless there because the credentials were already inert. That is false for the case this branch introduces: authenticateRouteToken and authenticateVmAuthorization both accept `paused`, so route tokens on a volume-backed machine stay valid after this write, for the rest of their 30-day lifetime. Keep the behavior, which is what getVm and the reconcile cron already do for the same observation, and replace the comment with what is true. Also drop the stale "next fleet refresh drops it" line, correct "seven call sites" to eight, stop asserting a resurrection path that is not implemented, and type usageEventSource as VmDestroySource so an unknown source cannot silently degrade in PostHog. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(cloud): reproduce lost destroy ledger retry * fix(cloud): retire missing provider compute durably * fix(cloud): durably drain observed destroy cleanup * fix(cloud): keep destroy cleanup queue fair and indexed * fix(vm): make observed cleanup retries index-safe * test(cloud): reproduce account deletion cleanup loss * fix(cloud): retain pending destroy cleanup on account deletion * fix(cloud): transfer account cleanup to durable outbox * fix(cloud): reject malformed cleanup outbox rows * fix(cloud): validate cleanup outbox shape * Persist explicit destroy cleanup for retry * fix cloud observed cleanup retries and concurrent migration * fix migration retry for invalid concurrent indexes --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
* feat: surface configured actions and launchers * fix(actions): preserve surface button action identity * fix(actions): report exact placements and global edit target * test(actions): pin exact surface placement identity * test(config): preserve surface action references * fix(actions): drop hidden surface action references * test(config): drop references for hidden surface buttons * fix: localize configured actions empty state * fix: sync upstream CmuxFoundation import repair (manaflow-ai#13236) Applies upstream 968cef2 to unblock this PR's app-host compilation. * fix(actions): localize discovery metadata labels * fix: localize actions discovery titles and open button Give the Actions menu item and dialog their own whole-string keys instead of borrowing the titlebar-layout debug label and concatenating, and label the open button with the existing Open cmux.json key rather than a raw path. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: document French identity for the actions discovery titles "Actions" is spelled the same in French, and cmux.json is a file name. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Stacked on #15221. GitHub will not let me base a pull request on a fork branch, so this diff contains #15221's commits as well. The only new commit is
abe6b75c9b8; review that one. Merging #15221 first will shrink this to just it.Why
Cmd-hovering a file path turns the cursor into a pointing hand. Cmd-hovering
#15221or a commit SHA gave you nothing, so the reference was clickable but invisible: the only way to find out was to click and see whether a browser opened. That is the worst of both, and it is the part of the brief about mouse-click affordances that actually bites.The awkward bit
A bare
#123only resolves against the pane's repository, and pointer motion runs on the main thread many times a second and cannot await an actor.So
GitHubRepositorySlugCachegrows a synchronous read.cachedSlug(forDirectory:)reports what a finished lookup left behind, through a lock-protected write-through mirror; the actor stays the source of truth and the mirror never starts a lookup of its own.The distinction that matters is
unresolvedversusresolved(nil). "Nobody has asked yet" and "this pane has no GitHub remote" must not look the same, or a hover arriving before the first lookup decides there is nothing to click, and that answer then gets memoized. Six tests cover the mirror, including expiry and both invalidation paths.An unresolved directory starts one background lookup, and that lookup re-runs the hover when it lands, so the affordance appears under a pointer that never moved.
Cost
Reading the visible line is a terminal text read, which is not something to do on every pointer event. The answer is memoized per terminal cell, since motion inside one cell cannot change it, and
pointerCell(at:)is that geometry split out ofvisibleWordPathSnapshotso a caller who only needs "did the cell change" does not pay for the read. Content scrolling under a stationary pointer does go stale; the next pointer move corrects it, which is the same tolerance the path hover already has.Ordering
A path wins where there is one, so a file named
#1still opens as a file. Only text that resolves to nothing on disk is offered to GitHub, matching what cmd-click does.Evidence, honestly
TerminalGitHubReferenceClickPolicytests in Cmd-click bare GitHub issue, PR, and commit references in the terminal #15221.Sources/.Follow-ups not in this PR
An underline under the reference, rather than only a cursor change, needs to draw into a surface Ghostty owns. That is a separate and much larger piece.
Changelog
Added: cmd-hovering a GitHub issue, pull request, or commit reference in terminal output now shows the same pointing-hand cursor as a file path.
🤖 Generated with Claude Code
Summary by cubic
Cmd-hovering a GitHub issue, pull request, or commit reference in terminal output now shows the same pointing-hand cursor as a file path, and cmd-clicking opens the reference on GitHub instead of doing nothing.
#123only resolves against the pane's repository, soGitHubRepositorySlugCachegains a synchronous read distinguishing "not resolved yet" from "no GitHub remote" so an early hover can't memoize a wrong answer.pointerCellgeometry helper to avoid a terminal text read on every pointer event; the memo is cleared on a path win, on cmd lift, when a lookup completion finds the hover no longer live, and a stale cell is recomputed rather than answered from a line that scrolled away.#1opens as a file..git/configand.git/HEADdirectly (the sandboxed runner'sgitis an xcrun shim) and use a distinctmanaflow-ai/cmux-reference-fixtureslug so resolving to the CI checkout can't satisfy them; negative assertions prove the pointer reached a cell.Written for commit 175e470. Summary will update on new commits.
Summary by CodeRabbit