Skip to content

Show the cmd-hover affordance over GitHub references - #15259

Open
teamleaderleo wants to merge 232 commits into
manaflow-ai:feat/clickable-linksfrom
teamleaderleo:feat/clickable-hover
Open

teamleaderleo wants to merge 232 commits into
manaflow-ai:feat/clickable-linksfrom
teamleaderleo:feat/clickable-hover

Conversation

@teamleaderleo

@teamleaderleo teamleaderleo commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Stacked on #15221. GitHub will not let me base a pull request on a fork branch, so this diff contains #15221's commits as well. The only new commit is abe6b75c9b8; review that one. Merging #15221 first will shrink this to just it.

Why

Cmd-hovering a file path turns the cursor into a pointing hand. Cmd-hovering #15221 or a commit SHA gave you nothing, so the reference was clickable but invisible: the only way to find out was to click and see whether a browser opened. That is the worst of both, and it is the part of the brief about mouse-click affordances that actually bites.

The awkward bit

A bare #123 only resolves against the pane's repository, and pointer motion runs on the main thread many times a second and cannot await an actor.

So GitHubRepositorySlugCache grows a synchronous read. cachedSlug(forDirectory:) reports what a finished lookup left behind, through a lock-protected write-through mirror; the actor stays the source of truth and the mirror never starts a lookup of its own.

The distinction that matters is unresolved versus resolved(nil). "Nobody has asked yet" and "this pane has no GitHub remote" must not look the same, or a hover arriving before the first lookup decides there is nothing to click, and that answer then gets memoized. Six tests cover the mirror, including expiry and both invalidation paths.

An unresolved directory starts one background lookup, and that lookup re-runs the hover when it lands, so the affordance appears under a pointer that never moved.

Cost

Reading the visible line is a terminal text read, which is not something to do on every pointer event. The answer is memoized per terminal cell, since motion inside one cell cannot change it, and pointerCell(at:) is that geometry split out of visibleWordPathSnapshot so a caller who only needs "did the cell change" does not pay for the read. Content scrolling under a stationary pointer does go stale; the next pointer move corrects it, which is the same tolerance the path hover already has.

Ordering

A path wins where there is one, so a file named #1 still opens as a file. Only text that resolves to nothing on disk is offered to GitHub, matching what cmd-click does.

Evidence, honestly

Follow-ups not in this PR

An underline under the reference, rather than only a cursor change, needs to draw into a surface Ghostty owns. That is a separate and much larger piece.

Changelog

Added: cmd-hovering a GitHub issue, pull request, or commit reference in terminal output now shows the same pointing-hand cursor as a file path.

🤖 Generated with Claude Code


Summary by cubic

Cmd-hovering a GitHub issue, pull request, or commit reference in terminal output now shows the same pointing-hand cursor as a file path, and cmd-clicking opens the reference on GitHub instead of doing nothing.

  • A bare #123 only resolves against the pane's repository, so GitHubRepositorySlugCache gains a synchronous read distinguishing "not resolved yet" from "no GitHub remote" so an early hover can't memoize a wrong answer.
  • An unresolved directory starts one background lookup that re-runs the hover when it lands, so the affordance appears under a pointer that never moved.
  • Answers are memoized per terminal cell via a new pointerCell geometry helper to avoid a terminal text read on every pointer event; the memo is cleared on a path win, on cmd lift, when a lookup completion finds the hover no longer live, and a stale cell is recomputed rather than answered from a line that scrolled away.
  • A file path still wins where one exists, so a file named #1 opens as a file.
  • The cmd-click harness asserts a slug-qualified ref lights on the first hover, a bare ref may take two (the first reads the pane's remote), and an ordinary word is checked repeatedly so a late lookup can't turn it into a link.
  • UI tests write the git fixture's .git/config and .git/HEAD directly (the sandboxed runner's git is an xcrun shim) and use a distinct manaflow-ai/cmux-reference-fixture slug so resolving to the CI checkout can't satisfy them; negative assertions prove the pointer reached a cell.

Written for commit 175e470. Summary will update on new commits.

Review in cubic

Summary by CodeRabbit

  • New Features
    • Command-click GitHub issue, pull request, and abbreviated commit references in the terminal to open them in a browser.
    • References can use the active repository or specify a repository directly. Hover feedback appears for recognized links, including after the active repository is resolved.

lawrencecchen and others added 6 commits September 27, 2026 22:41
Stack invitations carry a stored role applied on accept; reusable member-only
links redeem through an atomic claim; a per-team advisory lock keeps one admin.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Checkout, portal, subscription, and plan take an explicit teamId and require
team_admin. Requests without teamId keep the legacy implicit team.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Profile, emails and sign-in methods, notifications, sessions, API keys, and
account deletion now live under /dashboard/settings; /dashboard/team
redirects old hash links.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ance

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Cmd-hovering a file path turns the cursor into a pointing hand, so you know
before you click that there is something there. Cmd-hovering `manaflow-ai#15221` or a
commit SHA gave you nothing: the reference was clickable but invisible, which
is the worst of both, since you only find out by clicking and seeing whether a
browser opens.

The hard part is that a bare `manaflow-ai#123` only resolves against the pane's
repository, and pointer motion cannot await an actor. So the cache grows a
synchronous read: `cachedSlug(forDirectory:)` reports what a finished lookup
left behind, and says `unresolved` rather than "no repository" when nothing is
known yet. Confusing those two would mean a hover arriving before the first
lookup decides there is nothing to click, which would then be memoized.

An unresolved directory starts one background lookup and re-runs the hover
when it lands, so the affordance appears under a pointer that never moved.
The answer is memoized per terminal cell, because reading the visible line
costs a terminal text read and pointer motion inside one cell cannot change
it. `pointerCell(at:)` is that geometry split out of the snapshot so callers
that only need to know whether the cell changed do not pay for the read.

A path still wins where there is one, so a file named `#1` opens as a file.
Only text that resolves to nothing on disk is offered to GitHub, matching what
cmd-click already does.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository: manaflow-ai/cmux/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: 2b2adc6f-27c2-4294-a5ee-69bd0ffff3f9

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

The terminal detects GitHub issue, pull-request, and commit references. It uses explicit repository slugs or resolves a slug from the pane’s repository. Eligible references open through terminal link routing.

Changes

GitHub references in terminal

Layer / File(s) Summary
Reference detection and click policy
Packages/macOS/CmuxTerminalCore/Sources/CmuxTerminalCore/LinkRouting/TerminalGitHubReferenceDetector.swift, Packages/macOS/CmuxTerminalCore/Sources/CmuxTerminalCore/LinkRouting/TerminalGitHubReferenceClickPolicy.swift, Packages/macOS/CmuxTerminalCore/Tests/CmuxTerminalCoreTests/TerminalGitHubReference*Tests.swift
The detector recognizes issue, pull-request, and commit references, validates repository slugs, and builds GitHub URLs. The click policy opens references with sufficient repository information, requests resolution when needed, and ignores other clicks.
Repository slug lookup and cache
Packages/macOS/CmuxGit/Sources/CmuxGit/GitHubRepositorySlugCache.swift, Packages/macOS/CmuxGit/Tests/CmuxGitTests/GitHubRepositorySlugCacheTests.swift
The cache stores per-directory lookup results, shares concurrent lookups, expires entries, and supports synchronous reads. Tests cover cached misses, expiry, concurrent callers, and invalidation during lookup.
Terminal hover and Command-click integration
Sources/GhosttyTerminalView.swift, Sources/AppDelegate.swift, cmuxUITests/TerminalCmdClickUITests.swift, dogfood/scenarios/github-reference-click-tour.json
The terminal uses the detector and cache for GitHub-reference hover and Command-click handling. Fixtures, UI tests, and a dogfood scenario cover explicit and pane-derived repository slugs, commit references, and non-matching text.

Priority: ⬇️ Low

Estimated code review effort: 3 (Moderate) | ~25 minutes

Change: Feature

Sequence Diagram(s)

sequenceDiagram
  participant User
  participant GhosttyTerminalView
  participant TerminalGitHubReferenceClickPolicy
  participant TerminalGitHubReferenceDetector
  participant GitHubRepositorySlugCache
  participant TerminalLinkOpenCoordinator
  User->>GhosttyTerminalView: Command-click a terminal token
  GhosttyTerminalView->>TerminalGitHubReferenceClickPolicy: Evaluate runtime outcome and visible line
  TerminalGitHubReferenceClickPolicy->>TerminalGitHubReferenceDetector: Detect reference and repository requirement
  TerminalGitHubReferenceDetector-->>TerminalGitHubReferenceClickPolicy: Return reference or repository-resolution decision
  GhosttyTerminalView->>GitHubRepositorySlugCache: Resolve pane repository slug when needed
  GitHubRepositorySlugCache-->>GhosttyTerminalView: Return cached or discovered slug
  GhosttyTerminalView->>TerminalLinkOpenCoordinator: Open resolved GitHub URL
Loading

Suggested reviewers: austinywang

Merge Risk: 🔵 Low · up to 0eda9

The change is mergeable with owner awareness: scrolling or new output can leave the hover cursor stale, overlapping lookups can schedule redundant tasks, and the hover test may fail under slow CI conditions.

Security Architecture Review

Security architecture risk: 🔵 Low · up to 0eda9

Opening a reference still requires a user click, and generated links start at GitHub. No security bypass was established in the inspected flow. Repository freshness and the full scope of the stacked changes remain incompletely verified.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — An author of terminal output can influence the GitHub repository and reference selected by a deliberate user click. In the inspected path, this influence reaches browser navigation through the local application; it does not select an arbitrary initial URL host or become a Git command argument. Broader downstream navigation behavior was not audited.

Trust Boundaries and Controls

  • observed — The fallback rejects runtime-handled clicks and revalidates pane identities after asynchronous discovery. At the inspected head, the coordinator rejects remote-initiated requests before routing and applies its existing browser and destination policies. These head-side controls are not claimed to have been introduced by this PR.

Resilience and Maintainability Implications

  • observed — Hover memoization can retain an outdated affordance when terminal content changes within the same pointer cell. Clicking independently reads a snapshot and evaluates policy rather than consuming the hover result, so that memo is not an authorization token for URL opening.
  • observed — Repository answers remain reusable for 600 seconds, and delayed clicks retain click-time directory and text. These behaviors existed in the stacked parent; hover now permits earlier cache population. Freshness after remote changes and its effective exposure were not established by integration evidence.

Important

Pre-merge checks failed

Please resolve all errors before merging. Addressing warnings is optional.

❌ Failed checks (3 errors, 1 inconclusive)

Check name Status Explanation Resolution
Cmux Swift Concurrency ❌ Error The diff adds unowned fire-and-forget tasks in Sources/GhosttyTerminalView.swift. warmRepositorySlug(forDirectory:) starts Task { @MainActor ... } at lines 9064-9087 and discards its handle, alt… Store the hover lookup task on GhosttyNSView and cancel or replace it when command hover ends, the directory or target cell changes, or the view leaves its window. Clear the stored handle when the task finishes and check cancellation befo…
Cmux Architecture Rethink ❌ Error The new GitHubRepositorySlugCache adds a lock-backed side channel for state already owned by the actor. entries is stored in the actor, but Mirror.entries stores a second mutable copy behind `NS… Make one repository-slug owner authoritative. Remove the lock-backed Mirror, cachedSlug(forDirectory:), and the process-wide singleton access from the view. Route hover and click through an injected cache/coordinator that owns lookup, e…
Cmux No Test Or Debug Seam In Production Source ❌ Error The PR adds test-only observability to production Swift. Sources/GhosttyTerminalView.swift extends the existing debugSimulateCommandHoverDetails accessor with gitHubHoverActive and `gitHubHoverC… Remove the GitHub fields from debugSimulateCommandHoverDetails and remove the test-only joinedPendingLookupCount production member. Move test observation into the test target and use @testable import with the smallest necessary `priva…
Docstring Coverage ❓ Inconclusive Docstring coverage is 43.82% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 89 functions across 8 files. (3 skipped: … Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (21 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and concisely describes the main user-visible change: showing the Command-hover cursor affordance over GitHub references.
Description check ✅ Passed The description explains the problem, implementation, behavior, testing evidence, limitations, follow-ups, and changelog entry. It omits the template's Demo Video and Checklist sections, but the core …
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Cmux Cloud Persistent Session And Early Input ✅ Passed PASS: The reviewed diff adds GitHub reference detection, hover/click routing, cache logic, and UI fixtures. It does not change Cloud terminal creation, cmux-tui transport, manual renderer admission, r…
Cmux Swift Actor Isolation ✅ Passed No new actor-isolation failure matches the rule. GitHubRepositorySlugCache is an actor. Its shared mutable Mirror is explicitly @unchecked Sendable and protects its state with NSLock, with com…
Cmux Swift Blocking Runtime ✅ Passed The production diff adds no semaphores, blocking waits, sleeps, delayed dispatch, polling loops, or main-queue synchronous dispatch. It adds a small NSLock-protected mirror in GitHubRepositorySlugCach…
Cmux Browser Automation Off-Main ✅ Passed PASS: The PR does not change browser socket automation. The review-scoped diff leaves Sources/TerminalController.swift and ControlCommandExecutionPolicy.swift unchanged. The added code handles ter…
Cmux Expensive Synchronous Load ✅ Passed No prohibited synchronous agent-history load is introduced. The new interactive hover path calls only GitHubRepositorySlugCache.cachedSlug(forDirectory:), which reads a lock-protected in-memory mirr…
Cmux Cache Substitution Correctness ✅ Passed PASS: The new cache is used only for transient terminal hover/click resolution. GitHubRepositorySlugCache has no persistence, history, undo, or snapshot consumer. GhosttyTerminalView uses `cachedS…
Cmux No Hacky Sleeps ✅ Passed PASS. The authoritative diff contains only Swift source/tests/UI tests plus one JSON dogfood scenario. It contains no changed TypeScript, JavaScript, shell, or non-Swift build/runtime script. The fixe…
Cmux Algorithmic Complexity ✅ Passed PASS: The changed production code does not introduce a prohibited complexity pattern. GitHubRepositorySlugCache uses dictionary lookups and a write-through dictionary mirror. `TerminalGitHubReferenc…
Cmux Swift @Concurrent ✅ Passed PASS. The changed production async work uses valid isolation boundaries. GitHubRepositorySlugCache.slug(forDirectory:) is actor-isolated and accesses actor state, so it must not be @concurrent; UI…
Cmux Swift Package Boundaries ✅ Passed The reusable GitHub reference logic is package-bound. GitHubRepositorySlugCache is in the CmuxGit SwiftPM target, while TerminalGitHubReferenceDetector and TerminalGitHubReferenceClickPolicy a…
Cmux Swiftpm Lockfiles ✅ Passed PASS. The PR changes Swift source and tests in Packages/macOS/CmuxGit and Packages/macOS/CmuxTerminalCore, but it does not change any Package.swift, Package.resolved, .gitignore, workflow, o…
Cmux Swift Logging ✅ Passed The production diff adds no unguarded print, debugPrint, dump, NSLog, or ad hoc diagnostic file logging. The only new production diagnostic is the cmuxDebugLog call in openGitHubReference, guarded…
Cmux User-Facing Error Privacy ✅ Passed PASS. The production path adds GitHub reference detection, cursor state, and user-initiated link opening. It adds no user-facing error, alert, recovery text, API error body, or command error output. T…
Cmux Full Internationalization ✅ Passed The PR adds no production user-facing copy. The Swift changes implement cursor behavior, reference parsing, cache logic, and GitHub URL routing. The only displayed text addition is #847 or a supplie…
Cmux Swiftui State Layout ✅ Passed The PR does not introduce a SwiftUI state or layout pattern covered by the rule. The changed hover state and async task are private members and event handling in GhosttyNSView: NSView, an allowed Ap…
Cmux Swift Auxiliary Window Close Shortcuts ✅ Passed PASS: The PR adds GitHub reference detection, hover, click routing, and test fixtures. The authoritative diff adds no user-visible NSWindow, NSPanel, NSWindowController, SwiftUI Window, or WindowGroup…
Cmux Source Artifacts ✅ Passed All 11 changed paths are hand-written Swift source, unit/UI tests, or the deliberate dogfood/scenarios/github-reference-click-tour.json test fixture. The authoritative diff contains no binary files,…
Full details: Docstring Coverage

Explanation

Docstring coverage is 43.82% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 89 functions across 8 files. (3 skipped: 1 unsupported, 2 too large.)

Full details: Cmux Swift Concurrency

Explanation

The diff adds unowned fire-and-forget tasks in Sources/GhosttyTerminalView.swift. warmRepositorySlug(forDirectory:) starts Task { @MainActor ... } at lines 9064-9087 and discards its handle, although the task performs a Git lookup and later changes the cursor state. The task can outlive the command hover and view lifecycle. attemptGitHubReferenceOpen adds a second discarded task at lines 9140-9151 for delayed link opening. This matches the modernization rule for fire-and-forget tasks with meaningful lifecycle. The new GitHubRepositorySlugCache actor itself uses async/await appropriately, and the added test continuations are allowed test synchronization.

Resolution

Store the hover lookup task on GhosttyNSView and cancel or replace it when command hover ends, the directory or target cell changes, or the view leaves its window. Clear the stored handle when the task finishes and check cancellation before updating cursor state. Manage the delayed click lookup through a caller-owned, cancellable task or operation, and cancel it when the captured pane or view is invalidated. Keep the existing identity checks as an additional stale-result guard.

Full details: Cmux Architecture Rethink

Explanation

The new GitHubRepositorySlugCache adds a lock-backed side channel for state already owned by the actor. entries is stored in the actor, but Mirror.entries stores a second mutable copy behind NSLock; slug(forDirectory:) writes the actor and mirror separately, while cachedSlug(forDirectory:) reads only the mirror. removeAll() must coordinate both stores. The process-wide shared instance extends this duplicate ownership across all terminal views. This violates the rule's prohibition on new caches, singletons, side channels, and locks that leave shared state representable in multiple forms. The documented actor-as-source-of-truth comment does not remove the second state representation or its stale-read and invalidation-race class of bugs.

Resolution

Make one repository-slug owner authoritative. Remove the lock-backed Mirror, cachedSlug(forDirectory:), and the process-wide singleton access from the view. Route hover and click through an injected cache/coordinator that owns lookup, expiry, in-flight requests, and invalidation. Deliver an immutable repository-slug value snapshot with its generation to the main-actor terminal hover coordinator, and accept it only when the directory and generation still match. Use that same coordinator for both hover and click. The first migration cut is to replace Mirror with this read-only snapshot boundary and add a test proving that invalidation and an older lookup cannot produce an accepted snapshot.

Full details: Cmux No Test Or Debug Seam In Production Source

Explanation

The PR adds test-only observability to production Swift. Sources/GhosttyTerminalView.swift extends the existing debugSimulateCommandHoverDetails accessor with gitHubHoverActive and gitHubHoverCell; cmuxUITests/TerminalCmdClickUITests.swift reads these fields to validate hover behavior. This worsens an existing debug seam. Packages/macOS/CmuxGit/Sources/CmuxGit/GitHubRepositorySlugCache.swift also adds joinedPendingLookupCount, documented and referenced only by tests, to expose internal lookup behavior. The changes therefore introduce test/debug seams in production Sources/ code.

Resolution

Remove the GitHub fields from debugSimulateCommandHoverDetails and remove the test-only joinedPendingLookupCount production member. Move test observation into the test target and use @testable import with the smallest necessary private-to-internal widening, or isolate a genuinely required debug facility in a dedicated debug file or folder. Follow the reference fix in #6452.

✨ Finishing Touches 💡 1
🧪 Generate unit tests (beta)
  • Create a new PR
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

All contributors have signed the CLA ✍️ ✅
Posted by the CLA Assistant Lite bot.

Review of the hover commit found the memo answering for a cell the pointer had
already left. `||` short circuits, so a cell that resolves to a file path never
reaches the GitHub branch, and the memo is a single slot invalidated only by
being overwritten. Cross a path-bearing cell, let the terminal scroll, come
back, and the cursor uses the answer from a line that is no longer there: no
pointing hand over a reference that a cmd-click would open, or a pointing hand
over prose. Clearing the memo when a path wins keeps the saving the short
circuit exists for.

The background repository lookup had the same class of problem from the other
side. Its completion is the only asynchronous way back into the hover, and it
checked the modifier state and nothing else. `updateWordPathHover` falls back
to the last in-bounds point it saw, so a lookup landing after the pointer left
the view, or after the pane was closed, would push a pointing hand with nothing
left to pop it. It also dropped the selection check every synchronous caller
makes, so a lookup landing mid cmd-drag lit the affordance over a selection.
The completion now proves it is still a live hover before recomputing.

The memo's own comment claimed the next pointer move corrects staleness. Motion
inside a single cell does not, so the comment says what the code does. The
hover debug seam now also reports the GitHub answer and the cell it was
computed for, which is what a tour needs to walk path cell to reference cell
and back and see the memo follow.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Review of abe6b75c9b8 by a review subagent, correctness first. Fixed in 64445a9c22c.

Review

  1. The memo answers for a cell the pointer already left. hasTarget = resolution != nil || gitHubReferenceIsUnderPointer(...) short circuits, so a cell that resolves to a file path never reaches the GitHub branch, and gitHubHoverCell is a single slot invalidated only by being overwritten. Cross a path-bearing cell, let the terminal scroll, come back: the cursor uses the answer from a line that is gone. Deterministic, and in the feature this PR exists to add.
  2. The background lookup's completion checks NSEvent.modifierFlags and nothing else. updateWordPathHover falls back to the last in-bounds point it saw, so a lookup landing after the pointer left the view pushes a pointing hand over another app. If the pane was closed in the meantime, viewDidMoveToWindow has already run the only pop that was coming, and deinit does not pop, so the cursor stack is left one pointingHand deep with no owner.
  3. The same completion calls updateWordPathHover(cmdHeld: true) with the default suppressPathHover: false, dropping the selection check every other caller makes. A lookup landing mid cmd-drag lights the affordance over a selection, which is the regression that plumbing exists to prevent.
  4. gitHubHoverCell's comment says the next pointer move corrects staleness. Motion inside one cell does not.
  5. The doc comment for attemptGitHubReferenceOpen had been silently re-parented onto gitHubReferenceIsUnderPointer, so the hover function carried prose about click routing and the click function was undocumented.
  6. No test covers the view side at all: 1, 2 and 3 would have shipped green.
  7. Minor, low: a slug() call that joins an in-flight lookup can return before the mirror is written, costing one wasted task and actor hop. The memo is not keyed to pane identity, where attemptGitHubReferenceOpen guards on panel and workspace id. The hover now does a third terminal text read per cell change that the path resolution just did.

Explicitly attacked and found sound: the synchronous mirror (both write sites are paired, the generation guard covers removeAll() mid-flight, the nonisolated read of Sendable actor lets is legal under SE-0306), the pointerCell(at:) split (arithmetic-identical to the parent, checked line by line), and re-entrancy. No infinite loop exists on any branch: resolved(nil) terminates, an expired entry cannot re-arm inside the window, and the re-warm goes through a Task so there is no synchronous recursion.

Fixed

1, 2, 3, 4, 5. The memo is cleared when a path wins, so the short circuit keeps its saving without keeping a stale answer. The lookup completion now proves it is still a live hover (cmd held, pointer still resolvable inside this view) and passes shouldSuppressCommandPathHover(for:) like every synchronous caller. Comment corrected, doc comment moved back.

Partly 6: debugSimulateCommandHoverDetails now reports gitHubHoverActive and the memo cell, which is the seam a tour needs to walk path cell to reference cell and back.

Left

The view-level regression test itself. The seam is in place but the test is not written: the feature's async half gates on hardware modifier state that debugSimulateCommandHover cannot produce (it synthesizes [.command] with no key down), so a tour can cover findings 1 and 4 and cannot cover 2 or 3. I am not going to claim coverage I do not have.

The three low findings in 7. The wasted lookup round is one task and one actor hop, bounded. The pane-identity asymmetry costs at worst a lookup against a directory nobody is hovering. The extra text read is once per cell change, against a path hover that pays the same cost on every event.

No local build: Sources/ does not compile here and per the brief builds run in CI. The edited file parses (swiftc -parse); the compile verdict is CI's.

lawrencecchen and others added 7 commits September 28, 2026 03:49
Adds the typed route tree, session endpoint, flat search params, and the
shell, account menu, and team scope ported off next/navigation. Section
routes are placeholders until each section is ported.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Billing, TestFlight, CodeRouter, Cloud, Mobile devices, Settings, Teams,
Vault, home, and legacy redirects now run as typed TanStack routes with
TanStack Query data. Server-rendered reads move behind JSON endpoints:
/api/dashboard/{billing,coderouter}, /api/teams/[teamId]/billing,
/api/testflight (GET), /api/vm/access-grants, /api/vault/summary,
/api/vault/cli/auth/client, and /api/vault/sessions/[id]/head.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Next re-asserts the URL from a useInsertionEffect after every external
history write; TanStack's history treated that echo as a navigation and
scheduled a React update inside the insertion effect. Next's __NA writes
now skip TanStack's subscribers, and the history patch is removed when the
SPA unmounts. CodeRouter writes use useMutation, route titles come from
typed staticData, and new strings are translated for every locale.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The debug hover payload already reported the GitHub half of the answer,
but nothing read it. A scripted tour cannot: its hover step holds no
modifiers, and this affordance only exists while Command is down.

The harness does hold it, so assert the three cases directly. The
slug-qualified form has to light on the first hover because it needs no
lookup. The bare form is allowed to take a hover or two, since the first
one is what starts reading the pane's remote, so that test polls. The
ordinary word is checked repeatedly rather than once, because the failure
worth catching is a late lookup turning a plain word into a link after
the fact.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review points, same class as the ones on the click tests.

The ordinary-word hover test asserted only that the affordance stayed off,
which is also true when the pointer missed the token, when the surface is
gone, or when the feature is deleted. Each hover now also asserts the
harness reported `gitHubHoverCell`, which is set the moment the pointer
resolves to a cell and before any decision about what is in it.

Both fixtures claimed `manaflow-ai/cmux`, the repository CI runs from, so
resolving to the checkout instead of the fixture directory would have
looked the same. They now claim `manaflow-ai/cmux-reference-fixture`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Review: cmd-hover test coverage (cd65ca5a70d, 5b947a57660)

A review subagent went over the three added hover tests. Nothing blocked
compilation, and it confirmed the payload contract the assertions rely on:
Sources/AppDelegate.swift:3248 nests the hover dictionary under
lastCommandResult, and Sources/GhosttyTerminalView.swift:9238 writes
gitHubHoverActive as the string "1" or "0".

Fixed, in 6e4a4208432:

  • testCmdHoverOverAnOrdinaryWordShowsNoAffordance asserted only that the
    affordance stayed off. That is also true when the pointer missed the token
    entirely, and it would stay green with the affordance deleted. Each hover now
    also asserts the harness reported gitHubHoverCell, which
    gitHubReferenceIsUnderPointer sets the moment the pointer resolves to a
    cell, before any decision about what is in it. So the test now proves it was
    answering about something.
  • Both hover fixtures claimed manaflow-ai/cmux, the repository CI runs from.
    If a pane had resolved to the checkout instead of the fixture directory, the
    result would have looked the same. They now claim
    manaflow-ai/cmux-reference-fixture.

Left:

  • Thread.sleep in the two polling loops, where the rest of the file uses
    waitForCondition, which pumps the runloop. The review called it style rather
    than risk and I agree: worst case the bare-reference loop reports failure at
    about 126 s, and in practice runCommand throws at 10 s first. Not worth
    churning the tests for.
  • The three hover tests sit after two private helpers, breaking the file's
    tests-then-helpers grouping. Leaving it; moving them would make the diff
    harder to read against the branch below this one.

Also worth noting for this PR: the review points out the negative test now
covers the false-positive direction properly, but the late-lookup direction is
covered on #15221 instead (a bare #847 whose remote is not GitHub). No hover
equivalent was added, because the hover path memoizes per cell and the click
test already pins the policy decision.

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

CI failure attribution

CI failed on 8387d20445 (run 36714654335 attempt 1): 1 unknown.

Job Verdict Why
ui-tests unknown no known signature; failed step: Refuse a fork

Not re-run automatically: ui-tests is not a machine failure.

Written by scripts/ci/classify_failures.py (ci-failure-attribution.yml); signatures are its SIGNATURES table. A machine verdict is the runner's fault, not this PR's.

@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Not mergeable right now, and it is not this branch.

CI / ci-status (required) is red because
CI / guards / workflow-guard-tests / app-host-execution fails with:

AssertionError: [] is not true : agent-activity-reorder.json has no paths globs

That guard is tests/test_ci_pr_media.py:75, added to main in the same wave as
dogfood/scenarios/agent-activity-reorder.json, which has only launch and
steps and no paths key. So main requires every scenario to carry paths globs
and ships one that does not. Every PR against main picks it up through the merge
ref.

Everything else red here cascades from it: the Linux gate declines macOS, so
macOS compile admission stops at "Stop before compiling after a declined fast
Linux gate" and this branch's macOS code was never compiled at all. Same eight
failures, same order, on #15221, #15259 and #15325.

#15360 ("dogfood: give the agent activity reorder tour paths globs") already
fixes it and is mergeable, so I am not duplicating it. Waiting on that to land,
then merging main in and re-running.

@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

This branch was about to revert the tour fix, and now it does not

Checking why CI was red here, I compared this branch against #15221 and found a
worse problem than the red checks.

The tour scenario here was the pre-rewrite copy.
dogfood/scenarios/github-reference-click-tour.json at dbd4660a78b was
byte-identical (sha1 8cdbc0c887) to the version at 09b9ed6584a, before I
rewrote it on #15221. The rewrite is what gives each pane a real .git fixture
so the panes sit in a repository, adds the top-level paths globs, and makes
every after-click capture screen-wide. This branch carried the old 36-step,
no-paths copy.

Both PRs are based on main and both touch that file, so whichever merged
second would have won. If this one had merged after #15221, it would have
quietly put the broken tour back. Nobody would have seen it in review: the file
shows as unchanged against main on one PR and as a large diff on the other,
and neither view says "this reverts the other".

The cause is that the two earlier merges of feat/clickable-links into this
branch (48008ccb9fb, 5b947a57660) both predate the rewrite, and the later
merge of origin/main could not bring it in because it is not on main yet.

Fixed by merging the current feat/clickable-links head
(ca0240bbced) in: bfd4964e876. The tour file here is now the rewritten one
(sha1 865fae8e72, 37 steps, 5 paths entries, valid JSON), and
TerminalCmdClickUITests merged without duplication: 5 cmd-click tests from
#15221 and 3 hover tests from here, all present once.

New head: bfd4964e8766cadfca647b5f24fc39aa26301150. The queued fleet
dogfood request has been repointed at it; the older SHA should not be built.

On the diff you are reading

This PR is not stacked on #15221 in GitHub's eyes, because #15221's head branch
lives on the fork and a PR into this repository cannot take a fork branch as its
base. So the diff below still contains all of #15221's cmd-click work. Only the
hover half is new here:

 GitHubRepositorySlugCache.swift        |  63 +++++-
 GitHubRepositorySlugCacheTests.swift   |  84 ++++++++
 Sources/GhosttyTerminalView.swift      | 169 +++++++++++++++-
 cmuxUITests/TerminalCmdClickUITests.swift | 186 ++++++++-------

Once #15221 lands, the rest drops out on its own.

The red checks were not this branch

workflow-guard-tests / ci and .../app-host-execution failed, and the
cascade below them (linux-preflight, the macOS admission gate, ci-status)
follows from that, not from anything here.

The diff touches no scripts/, no tests/, no .github/. On main at
7171ea8c8bb both of those jobs are green. Locally on this head,
tests/test_ci_owned_spm_scratch.py passes 5 runs out of 5 and
tests/test_ci_ui_tests_dispatch.py 3 out of 3. The two attempts here failed on
different tests in the same files, and the annotations carry an HTTP 403 and a
gh api exit 124 (a timeout), which is what the failures look like when the
suite cannot reach GitHub rather than when it has found a bug.

So: infrastructure flake, re-running. I am not claiming this branch is green
until it actually is.

@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Correction: half of that red was this branch, and I said it was not

An hour ago I wrote here that the guard failures "were not this branch" and
called them infrastructure flake. That is wrong for one of the two jobs, and I
had the means to check it before I posted.

workflow-guard-tests / app-host-execution was failing on my diff. That
group runs tests/test_ci_pr_media.py, and
test_checked_in_tours_are_valid_and_the_default_exists asserts that every
checked-in dogfood scenario declares paths globs. The stale tour file this
branch was carrying had none. Reconstructing the dbd4660a78b tree and running
the test against it:

AssertionError: [] is not true : github-reference-click-tour.json has no paths globs

So the repository already had a guard for exactly the reversion I described in
the comment above, it fired, and I read past it. Against the merged head
bfd4964e876 the same test run reports 0 failures.

What I did wrong is specific: I ran the two suites I had guessed were involved
(test_ci_owned_spm_scratch.py, test_ci_ui_tests_dispatch.py), saw them pass,
and generalised from that to the whole job. I never ran the suite that was
actually failing, because I never looked up which step in
app-host-execution had failed. The annotation text I did read
("No ci-ui-tests.yml run titled ... appeared") is ordinary stdout from an
earlier passing test in the same step, and I took it for the failure.

workflow-guard-tests / ci I still cannot explain. That is a different
step, tests/test_ci_owned_spm_scratch.py, it named a different test on each of
the two attempts, and it passes 5 out of 5 locally and is green on main at
7171ea8c8bb. I am not calling that flake again until the new head has run: if
it comes back red, it is mine and I will chase it.

The re-run is on bfd4964e8766cadfca647b5f24fc39aa26301150.

One thing worth keeping: the near-miss in the comment above needs no new guard.
test_ci_pr_media.py already catches a scenario that loses its paths, which
is the tripwire a silent revert of a tour file trips. It worked. The gap was me.

@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Out-of-band UI test run at bfd4964e876

The ui-tests job cannot run here: this PR is from a fork and changes
cmuxUITests/TerminalCmdClickUITests, which trips the unconditional Refuse a fork step in ci.yml. So I dispatched the class directly on an owned runner.

Run: https://github.com/manaflow-ai/cmux/actions/runs/36433190400

All eight reference tests pass, the three hover ones included:

passed testCmdHoverOverABareReferenceShowsTheAffordanceOnceTheRepositoryIsKnown
passed testCmdHoverOverASlugQualifiedReferenceShowsTheAffordance
passed testCmdHoverOverAnOrdinaryWordShowsNoAffordance
passed testStationaryCmdClickBareIssueReferenceOpensItAgainstThePaneRepository
passed testStationaryCmdClickSlugQualifiedReferenceNeedsNoRepository
passed testStationaryCmdClickCommitShaOpensThatCommit
passed testStationaryCmdClickOrdinaryWordOpensNothingInARepository
passed testStationaryCmdClickBareReferenceOpensNothingWhenTheRemoteIsNotGitHub

Thirteen other tests in the class failed: the spaced-path and alt-screen path
cases, the two Markdown trailing-period cases, and
testSSHCmdClickDownloadsSpacedPathIntoPreview, which fails with
Error Domain=SSHPreviewFixture Code=1 before it clicks anything.

I am not calling those pre-existing yet, because I have been wrong about that
before on this stack. What I can say now: the identical thirteen failed on
#15221's head (run 36433313040), which is a different commit, and none of them
is a test this stack adds. I have the same class dispatched at main
(bc28bc4c31a) and will post the comparison here either way.

A hover tour cannot show this feature, so I am not writing one

I was planning to move the cmd-hover steps out of #15221's tour and into a tour
here, since this is the branch where gitHubReferenceIsUnderPointer actually
exists. That plan is wrong and I am dropping it.

The affordance this PR adds is a cursor shape. The tour's shot step calls
window.screenshot(), or XCUIScreen.main.screenshot() for a screen-scoped
shot, and neither draws the mouse cursor into the image. A hover frame would
therefore be byte-comparable to the plain terminal frame before it, and would
read as "the affordance did not appear" no matter how well the feature works.

The three hover XCTests above are the evidence for this behavior. They assert
on the resolved hover state rather than on pixels, which is why they can see an
affordance a screenshot cannot.

austinywang and others added 24 commits September 30, 2026 18:59
* test: cover descriptor-backed artifact thumbnails

* fix: decode artifact thumbnails from verified descriptors

---------

Co-authored-by: Leo Li <cheerleaderleo@outlook.com>
…s at once (manaflow-ai#15869)

* Revoke a removed member's team VM access on the Stack webhook

Stack team_membership.deleted now detaches every tunnel of that user from
the team network and drops their identity snapshot, so open terminals and
browsers lose the route at once instead of after the 10 minute cron.
Tunnel enrollment reconciles against the complete, fresh team list and
never detaches when that list is incomplete, so a Mac keeps every team
network it belongs to.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep Cloud surfaces from several teams open across team switches

Each Cloud provider, workspace binding and browser panel records its owning
team, and every VM request for a live surface sends that team instead of the
active one. A team switch now only rescopes the sidebar and new machines;
open terminals and browsers of other teams stay connected, including after
restore. A 404 vm_not_found, 403 or vm_owner_mismatch is permanent: the pane
shows that access was lost instead of freezing on its last frame.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: temporary hang diagnostics for Cloud graph suites

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
)

* Add dictation accessibility regression tests

Cover what third-party dictation tools need from the terminal's
accessibility element: AXValue shows the screen so a tool can confirm its
insertion (manaflow-ai#722), setting AXSelectedText types at the cursor (manaflow-ai#4953), a
multi-line value arrives as one bracketed paste, and an unbound Command
chord posted by another process does not type into the terminal (manaflow-ai#4153).

These fail before the fix that follows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Let third-party dictation tools insert into terminals

Dictation tools that insert through accessibility could not confirm their
insertion, could not use AXSelectedText, and a multi-line dictation ran
each line as its own command.

- AXValue now returns the active screen from a 500 ms snapshot, with
  NumberOfCharacters, VisibleCharacterRange, Line(for:) and String(for:)
  getters, and posts valueChanged after an AX insertion (manaflow-ai#722).
- setAccessibilitySelectedText commits at the cursor, and both setters are
  reported as settable (manaflow-ai#4953). A client that writes back the value it read
  with its text spliced in gets only its text typed, not the whole screen.
- Text with an interior line break goes through the paste path, so a
  bracketed-paste-aware shell or agent gets one block. Single-line text,
  including a trailing newline sent as Return, keeps typed semantics.
- A Command chord that misses the menu and every Ghostty binding is dropped
  when another process posted it, so a dictation hotkey such as
  Cmd+Option+C no longer types a stray character (manaflow-ai#4153).

The accessibility overrides move to GhosttyNSView+Accessibility.swift.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Harden the dictation AX splice and narrow the foreign chord drop

Review follow-ups:
- Diff an AXValue write against the last 8 distinct values handed to AX
  clients, so a tool that writes back an older read after the screen
  changed still gets only its text typed. Short values need a pure
  insertion, so a literal isn't trimmed to match a two-character prompt.
- Take the multi-line paste path only on a live surface, where the paste
  and the trailing Return share one input sequence, and end any IME
  preedit first.
- Drop only Command+Option chords posted by another process, the
  push-to-talk hotkey case, so remote-control and automation tools keep
  their other unbound Command chords.
- Cover a stale read, a multi-line value with a trailing newline, and
  assert the keyboard chord in the foreign-chord test is local.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover delayed accessibility edits

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: retain accessibility reads for delayed edits

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover accessibility history retention

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: an exact older AX read wins over a partial newer match

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: prefer an exact AX read over a partial newer match

With 30 s of history, a write spliced into an older read could match a
newer screen that shares most of its text and paste the older screen's
differing tail. Look for a pure insertion across the whole history first.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: announce AX inserts after terminal screen updates

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…anaflow-ai#10277)

* perf: read listening ports from the kernel instead of spawning lsof

The batched port scan ran `lsof -nP -a -p <pids> -iTCP -sTCP:LISTEN -Fpn`
every two seconds. Before lsof answers it calls close() on every descriptor
number up to `kern.maxfilesperproc`, which is 138,240 on current macOS, and
it forks a child that repeats the sweep. That is far more work than the
lookup itself, and it repeats for the life of the app.

`ListeningPortLookup` asks the kernel directly through libproc
(PROC_PIDLISTFDS + PROC_PIDFDSOCKETINFO) and keeps only TCP sockets in
TSI_S_LISTEN. Measured against lsof over 483 PIDs the results are identical,
and one scan drops from 200ms to under 7ms.

Completeness semantics are unchanged: a PID we may not read is a miss only
when its identity is also unreadable while it is still present, so a panel
behind the root `login` process can still retire its ports.

* perf: stop local port scans when the ports detail is hidden

Remote port scanning already follows `sidebar.showPorts` (issue manaflow-ai#6123), but
the local scanner kept running its two-second sweep no matter what. Nothing
displays or reports the result while the detail is hidden, so the work is
pure cost.

`SidebarWorkspaceDetailDefaults.portScanningEnabled` now holds the one rule
both paths read, and `TabManager` pushes changes to the local scanner
alongside the remote sessions it already updates.

Note this also blanks `listeningPorts` for local workspaces in the control
socket and CLI summaries while ports are hidden, which is the same trade
remote workspaces already make.

* Stop and restart the local port scan cleanly with the ports detail

Hiding the ports detail now also invalidates a burst that is already
running, the same way unregistering the last panel does, so its remaining
timers never scan. Showing it again rescans every registered panel once,
since ports that opened or closed while hidden were never seen and an idle
panel would otherwise keep a stale list until its next command.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Address review: clear hidden ports, move publication tests off lsof

- Hiding the ports detail now publishes no ports for every tracked panel
  and agent workspace, so the socket, CLI, custom sidebars and the command
  palette (which read published ports without the sidebar's visibility
  gate) don't keep a list frozen at the moment scanning stopped. A fresh
  agent request ID fences off any scan still in flight.
- A queued follow-up panel or agent scan no longer runs after scanning is
  turned off; a dropped agent request clears its in-flight mark.
- PortScannerPublicationTests fed ports through a fake lsof, which the
  scanner no longer calls. Both runners now supply ports through
  listeningPortsProvider, and the churn test counts per-PID lookups instead
  of lsof argv lists.
- Test for the cleared list; normalize the pbxproj; drop stale lsof
  wording from two doc comments.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Pass sets to the snapshot reconcilers when clearing hidden ports

Fixes the CI compile error in the previous commit: remove(keys:) takes a
Set, not an Array.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Leo Li <cheerleaderleo@outlook.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(cloud): cover idle hover accessory hit testing

* fix(cloud): keep idle row clicks available

* fix(cloud): hide stale row hover controls

* fix(cloud): keep unavailable display actions inert

* test(cloud): validate idle row hit target coordinates
* test: cover workspace group anchor behavior

* fix: route workspace group headers safely

* test: preserve generated group header fixtures

* fix: preserve generated group anchor state

* Update group anchor lifecycle fixtures

* Align group reorder fixtures with anchor semantics

* Fix remaining group fixture compile error

* Preserve generated-anchor reorder fixture coverage

* cmux: follow current Bonsplit tab ID API

* fix: close generated anchors consistently

* fix: scope generated anchor cleanup

* fix: protect confirmed groups during deletion

* ci: declare the vendor/bonsplit bump as forward

submodule-forward-only: allow vendor/bonsplit

The Bonsplit update is a forward move whose merge ancestry is hidden by the shallow submodule checkout.

* ci: preserve Bonsplit forward declaration after merge

submodule-forward-only: allow vendor/bonsplit

* test: cover confirmed group deletion cleanup

submodule-forward-only: allow vendor/bonsplit

Use the Bonsplit main commit that contains the existing tab ID API and current Liquid Glass changes.

* fix: preserve generated anchors with Dock content

* fix: retain current Bonsplit API compatibility
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…16150)

* ci: report a cancelled run instead of a routing failure

The `tests` gate in ci.yml and the `ios-tests` gate in test-ios.yml both
run with `if: ${{ always() }}` so that they always report, because both
are in the branch protection required list. A cancelled run therefore
reaches them with every job it depends on reporting `cancelled`, and
their routing check fails with only `changes: cancelled` or
`detect-ios-changes: cancelled` to show for it.

That message describes a routing defect, and the pull request does not
have one. It cost three pull requests a diagnosis each this week before
the answer turned out to be the same every time: rerun the cancelled
run. Reruns of two of them came back green with no change to the branch.

So say that instead. Each gate now has a step that runs only when the
run was cancelled and prints what happened and what clears it.

Both gates still fail on a cancelled run. A cancelled run tested
nothing, and reporting success would let a pull request merge on a green
required check that no suite stood behind. `cancelled()` is true only
for a cancelled run, so a job that timed out still reaches the routing
check and fails there, which is where a timeout belongs.

`cancelled()` is available in `jobs.<job_id>.if` and
`jobs.<job_id>.steps.if` only, not in a step's `env:`, so this is two
mutually exclusive steps rather than one step reading a variable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: detect the cancellation in the gate script, not in a step condition

The first version of this used a second step gated on `cancelled()`. That
only covers a cancelled workflow run: `cancelled()` says nothing about an
individual dependency, so a job cancelled on its own still fell through
to the routing verdict and still reported as a routing defect.

Both gates already read every dependency's result out of `toJSON(needs)`,
so the cancellation is visible there. Checking it in the script covers a
cancelled run and a cancelled dependency with one branch, needs no step
condition, and drops the second step.

The verdict is unchanged in both other cases: a run where everything
reported still exits 0, and a genuine routing failure still exits 1 with
the message it had before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…i#14868)

* Add setting actions, setting presets, and cmux config set

cmux.json actions can now change settings: "type": "setting" with a
path and one of set, toggle, cycle, or unset, and "type": "settingPreset"
to apply a named partial settings object from the new top-level
settingPresets. `cmux config get|set|unset|toggle|cycle|preset` drive the
same JSONConfigStore.apply path, so every entrypoint validates against
the schema, keeps comments and unrelated keys, and publishes one atomic
write under the cooperative writer lock.

Setting actions only run when the global cmux.json declares them; project
configs and packs can't ship a button that rewrites global settings.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Address review: empty preset objects, fail-closed trust, docs

- A preset's empty nested object merges nothing instead of replacing
  the whole section; a preset that sets nothing is refused.
- A setting action with no source path fails closed.
- Docs, schema, and comments say packs referenced by the global config
  may declare setting actions (they inherit its source path), and that
  confirm doesn't apply.
- Move the CLI help note below the subcommand list.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Honor confirm on setting actions and explain keys containing "."

A setting or settingPreset action with "confirm": true now asks before
saving and shows the equivalent `cmux config` command. The project-action
trust prompt never covers the global config, so this is the only prompt
these actions get.

A path that only resolves when one of its keys contains "." (for example
a workspaceGroups.byCwd entry for ~/src/app.web) is refused with a
keyContainsDot error that says so, instead of "isn't a cmux setting".
Such keys stay unsupported; there is no escaping syntax. Docs, the CLI
contract, and the cmux-settings skill say so.

Adds a test that packs the global config references may declare setting
actions while a project config's packs may not, and that confirm reaches
both the palette action and the tab bar button.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Write fractional setting values in their short form

JSONSerialization prints a Double with 17 significant digits, so
`cmux config set terminal.scrollSpeed 1.4`, a cycle entry, or a preset
leaf wrote 1.3999999999999999 into cmux.json and the CLI echoed it.
CmuxSettingValue now hands JSONSerialization an NSDecimalNumber built
from Swift's shortest round-trip text, and preset leaves are re-encoded
through CmuxSettingValue so every path writes 1.4.

CI run 36310670056 on e478696 caught this through the new
commandLineDescriptions and confirmationDialogShowsTheEquivalentCommand
tests; this adds a file-text assertion for set and preset.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Start setting changes from the live value; keep project configs out

Review fixes:

- Most settings live in UserDefaults, and cmux.json only overrides them.
  toggle, cycle, and get now read a key the file doesn't set from the
  app's UserDefaults (CmuxSettingLiveValues, backed by the setting
  catalog) before the schema default, so the first press after changing
  a setting in the Settings window flips what the user sees. The app
  reads its own defaults; the CLI reads the enclosing app's domain.
  `cmux config get` says where the value came from, and `unset` is
  described as removing the key from cmux.json.
- A project config could override a global setting action's title,
  shortcut, or confirm through an actions entry or a tab bar button.
  Both overrides are now ignored for setting actions unless they come
  from the global config.
- The failure alert falls back to a localized "couldn't read or save
  cmux.json" message for store errors that have no description.
- keyContainsDot only fires when the rejoined key exists in the file or
  looks like a path, so a typo under a map keyed by names stays an
  unknown path.
- `cmux config get --json` uses `path` and `file` like the other
  subcommands, plus `source`.
- The schema no longer says confirm doesn't apply to setting actions.
- The setting-action docs keys are translated into the 18 other web
  locales, and CHANGELOG has a line.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Only trust live values that are stored in cmux.json form

Second review pass:

- Some UserDefaults values aren't stored the way cmux.json spells them:
  app.minimalMode is a presentation-mode string, and
  app.keepWorkspaceOpenWhenClosingLastSurface is stored as the opposite
  flag. Lists and maps are often stored as text. The live resolver now
  maps those two explicitly and otherwise accepts only a scalar whose
  type the schema allows at that path (CmuxConfigSchemaPathLookup gains
  declaredTypes(at:)). A catalog-wide test fails if a key the resolver
  accepts stores a default that disagrees with the schema default, so a
  new transformed key has to get a mapping.
- Re-setting a fraction that's already in the file is no longer reported
  as a change: receipts and the no-op check compare numbers in the same
  short form they are written in.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Map the session content width sentinel back to false

The catalog test caught terminal.sessionContentMaxWidth: UserDefaults
stores -1 for "no cap", which cmux.json spells false. The live resolver
now maps widths under the minimum to false.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Move the changelog line to the end of Unreleased/Added

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix setting action docs rich text

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix setting presets from global packs

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix settings module import

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix config executor settings import

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add setting actions dogfood tour

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix setting actions tour JSON

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Refine setting actions dogfood tour

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Make setting actions tour reload config

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Isolate setting actions dogfood config

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Use isolated home in setting actions tour shell

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix setting actions tour cleanup

Use the isolated tour home when restoring config so cleanup cannot remove the user's config.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Harden setting action persistence

Prefer current global presets, guard symlink retargets during publication, and keep tour cleanup isolated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ow-ai#15217)

* test: predicted echo stays off for surfaces not known to be remote

A local shell that echoes slower than the threshold would draw predicted
glyphs, and every registered surface's output is copied and scanned while
the setting is on. Both tests fail on main.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Predict echo only on remote terminals; withdraw on pasted and sent input

Every terminal registered with the prediction center, and a local shell
was kept out only by its measured echo latency. A shell on a loaded Mac can
echo slower than the 25 ms threshold, and while the setting was on every
terminal's PTY output was copied and scanned on the main actor. Each
surface is now classified at its first keystroke after prediction starts
for it: only a terminal whose shell runs on another machine (an SSH or
Cloud machine, a remote tmux pane, another signed-in Mac) gets an active
engine, and the output inbox accepts only those surfaces, so a local
terminal's reads are dropped before any copy.

Input that bypassed the keystroke path (a paste completing, text and keys
sent over the socket, the TextBox, a paired phone) left an armed echo run
in place. A key typed right after a paste was drawn where the pasted text
was about to land. That input now withdraws the prediction, like an
editing key does.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Turn predictive local echo on by default and move it to Terminal settings

With prediction limited to terminals whose shell runs on another machine,
the setting leaves Beta Features and defaults on. It is now
Settings > Terminal > Predictive Local Echo and `terminal.predictiveLocalEcho`
in cmux.json and the schema. It stays stored under the former beta
UserDefaults key, so anyone who turned the beta off keeps it off.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Honor the catalog default for predicted echo; withdraw on remote tmux named keys

The center read its setting with UserDefaults.bool(forKey:), which reads an
unset key as off, so a catalog default of on would never reach it. It now
takes the default from the catalog. A transport-named key in a remote tmux
pane skipped the prediction path and left drawn glyphs in place.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Classify Dock terminals for predicted echo; reclassify when a remote session starts or ends

A remote terminal moved into a Dock was classified local because only the
workspace split tree was consulted, so it never predicted. Classification
now resolves the hosting Dock first, as terminal link opening does.

Classification was fixed at a surface's first keystroke, so a terminal typed
into before its remote session registered (a reconnect that keeps the
surface) never predicted in that runtime. A change to a workspace's remote
terminal set now withdraws and classifies the surface again at its next
keystroke.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: deleted text reappears after Backspace, Ctrl-U, Ctrl-W over in-flight predictions

A frame-composing harness (grid plus overlay every 16.7 ms) flags any column
the user deleted and saw blank that later shows text again. 6 of 8 fail.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep deleted predictions blank until the remote erases them

A Backspace over a glyph still in flight no longer lets its echo paint the
character back: the cell is drawn blank until the erase lands. Ctrl-U,
Ctrl-W and Option-Backspace mask in-flight glyphs the same way, and keys
whose effect is unknown (Return, arrows, pastes) raise a barrier behind
the glyphs already sent instead of dropping them, so nothing vanishes
and reappears. A glyph retyped onto a blank cell replaces the blank.

The simulation now checks blanks at the end of the run: a blank over text
already where it belongs fails only if the remote never rewrites that cell.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: Return answered by a newline suspends prediction; Ctrl-U after Return blanks the submitted line

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Resolve the barrier on the remote's newline; match Ctrl-U and Ctrl-W by character

Output answering Return, an arrow or a paste (CR LF, a redraw) resolved the
barrier only when printable; a newline counted as a misprediction, so fast
typing followed by Return could suspend prediction. A barrier that expires
unanswered guessed nothing and no longer counts either. Ctrl-U behind a
barrier acts on the new line and leaves the submitted glyphs drawn.
Ctrl-U and Ctrl-W are recognized by the layout's character rather than the
physical key.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Trace predicted echo in dev builds while a flag file exists

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: drain-before-parse draws echoed glyphs over the cells to their left

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: zsh highlighting suspends, images keep stale glyphs, slow jittery links stop drawing, main-actor scan cost

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: links over 1.5 s never arm; one stall stops drawing until the user pauses

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: type-ahead and masked password prompts draw secrets; status repaints stop prediction

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep predicted echo safe at password prompts, alive through rewrites and stalls

Found while dogfooding over a slow link and by adversarial tests:

- zsh answers the second key of a line with BS and a rewrite of the first
  ("\bec"), which withdrew everything as a misprediction and left prediction
  dark for the rest of the line. The engine now remembers what it saw echoed
  on the row and accepts a move back over it followed by the same text; CSI n D
  reports its count for this. zsh-syntax-highlighting recolouring is the same.
- A key typed ahead of a password prompt got its cooked-mode echo and armed
  the run, so the rest of the password was drawn. Drawing now needs two
  echoes in a row since the run last ended, and a `*` never counts.
- An overdue echo or unmodelled cursor motion now hides what is drawn but
  keeps matching the keystrokes in flight, so one stall no longer leaves the
  line untracked while the user keeps typing. The lifetime scales to three
  round trips on links slower than 500 ms.
- A cursor save and restore (a status line or clock repaint) is ignorable;
  kitty graphics and sixel images are disruptive.
- Output with nothing in flight is skimmed instead of classified byte by
  byte, and the inbox keeps at most 1 MiB per surface, reporting a drop.
- Backspace over an echo not yet painted blanks it until the erase. Blanks
  that output removes stay until the next presented frame, and the host
  drains a surface's pending output before retiring at a frame.

The breaker tests that armed on one echo now arm on two. The two render
race tests are disabled until ghostty tees after the parse.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: point the disabled render-race tests at the ghostty tee fix

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: read the final status outside the expectation

The determinism guard reads a duration literal inside an assertion as a
wall-clock check. This one is the engine's simulated clock.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: a predicted run over the layer-hosting terminal renders nothing

Ghostty makes the terminal view layer-hosting. The overlay drew in
draw(_:), which AppKit called with the host's whole bounds as the dirty
rect, and the run never reached the screen: on a dev build with a 1.8 s
link, ten predicted glyphs sat drawn, unhidden and on top for over a
second with nothing visible.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Render the predicted run into the overlay layer's contents

The terminal view is layer-hosting (ghostty's IOSurface layer is its
layer), and AppKit's draw pass for a subview of it never put the run on
screen. Render the run into a bitmap of the overlay's own size and set it
as the layer's contents from updateLayer, which does not depend on that
pass.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: a layout hold with no presented frame freezes the overlay

On a dev build the rendered-frame callback never fired: it is counted in
the Metal layer's nextDrawable, and ghostty presents through its own
IOSurface layer. The hold set when an erase removed blanks was never
released, so the overlay kept a blank over a cell the grid had already
redrawn, and hid every later prediction on that surface.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Release the layout hold when no frame arrives

The hold waits for a presented frame, and the host may never report one.
Release it after confirmationHold, the same bound a held confirmation
has, and include that deadline in nextExpiry so the host's timer re-anchors
the overlay.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add a predicted-echo recording script for dogfood GIFs

Puts a local sshd behind a delay proxy, opens a cmux ssh workspace on a
tagged build, types paced keys through the debug socket and records the
window as a GIF. The dogfood tour format cannot express a slow link, a
remote workspace, timed keys or video.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix predictive echo for consumed bindings

* pin Ghostty tee fix release

* fix predictive echo binding classification

* classify Ghostty local-only bindings

* fix Ghostty binding flag typing

* record local-only GhosttyKit checksum

* rebase Ghostty binding flag onto main

* pin Ghostty local-only binding fix

* record rebased GhosttyKit checksum

* rebase Ghostty binding fix onto current main

* fix Ghostty gitlink

* record current GhosttyKit checksum

* retain GhosttyKit checksum after main merge

* ci: record rebased GhosttyKit checksum

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ai#16140)

* test(ci): update-homebrew must keep triggering-run metadata out of scripts

update-homebrew.yml substitutes github.event.workflow_run.head_branch and
the dispatch input into its `run:` script as text. The test renders the
version step the way the runner does, with a hostile branch name, and
checks nothing executes; it also requires the gate to accept only the
real release workflow run from a tag push in this repository. Found by a
Codex Security audit (both models).

Red: python3 tests/test_ci_homebrew_untrusted_input.py

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(ci): update-homebrew takes run metadata through env and trusts only the real release run

Makes tests/test_ci_homebrew_untrusted_input.py green (previous commit).

Root cause: the version step substituted
github.event.workflow_run.head_branch and the dispatch input into its
script text, and the gate accepted any run of a workflow with the
release workflow's display name. The values now reach the script as
env vars (INPUT_VERSION, HEAD_BRANCH, EVENT_NAME), and the gate requires
the run to be .github/workflows/release.yml, from this repository, for a
push. Dispatch keeps working as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
manaflow-ai#16148)

* ci: survive an owned Mac whose Homebrew prefix the runner does not own

Run 36752535712 built the app-host and UI test product on
cmux-austin-mini-1-glaeda-1, 17 minutes, then lost the whole job in the
"Install tmux" step:

  Error: /opt/homebrew/Cellar is not writable. You should change the
  ownership and permissions of /opt/homebrew/Cellar back to your
  user account:
    sudo chown -R cmux /opt/homebrew/Cellar

brew refuses outright when the prefix belongs to another account, so
the step exits 1, "Resolve selectors against the built tests" is
skipped, and the job reports TEST_RESULT=failed with TEST_SUMMARY and
TEST_OUTPUT both empty. The dispatch looks like a test failure and names
no cause short of reading the raw log.

scripts/ci/brew-ensure.sh retries the install as the prefix owner, and
when even that cannot produce the command it fails with the runner name
and the package to provision. Both brew installs in the E2E action use
it. Each call site keeps its old inline path behind an -x check, because
this action runs against the tested revision's checkout, which may
predate the script.

This does not provision that mini; it stops one missing package from
throwing away a build that already succeeded, and makes the next
occurrence legible from the step summary.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* ci: reach the helper from the workflow revision and classify its failure

Review of the first commit found five ways the fix would not have helped the
run it was written for.

The helper was called workspace-relative, so it came from the tested revision.
test-e2e.yml checks the action out of the workflow's own revision into
.e2e-workflow with a sparse-checkout that did not list the helper, so a
re-dispatch of the ref from the failing run, and every step of a regression
bisect into older history, would have taken the old inline brew install and
failed identically. Add the helper to both sparse-checkout blocks and prefer
the .e2e-workflow copy at both call sites, the way the frames script already
does.

Guard on -f and invoke through bash instead of -x. A lost mode bit made the
ffmpeg branch fall through to an unconditional brew install, which fails on an
unowned prefix even when ffmpeg is already there: worse than no helper at all.

The retry now runs sudo -n, gated on sudo -n true, which is how the rest of
this action and run-in-console-session.sh treat passwordless sudo on these
Macs. Without -n, a host with a controlling tty would block on the prompt
until the job timed out, holding a Mac.

HOMEBREW_NO_AUTO_UPDATE was a command prefix on the first install only, and
sudo's env_reset would drop it regardless, so the retry could trigger a full
brew update inside a step that had already burned the build. Carry it over the
sudo hop with /usr/bin/env, as action.yml:545 does.

Every failure path now prints [cmux-ci machine: brew-provision], so
machine_failure.py classifies it directly instead of depending on Homebrew's
own wording reaching the log. A root-owned prefix gets its own message, since
brew refuses to run as root and no hop fixes that.

Verified: shellcheck and bash -n clean, actionlint clean on test-e2e.yml, both
YAML files parse, tests/test_ci_machine_failure.py passes with the new case (8
tests), tests/test_ci_self_hosted_guard.sh passes, and all eight branches of
the script exercised under brew/stat/sudo shims exit 0 only with the command
present and non-zero only with it missing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(ci): pin the new helper in the e2e sparse-checkout guard

test_the_build_runner_runs_the_tests_so_a_run_queues_once asserts the
sparse-checkout list verbatim, so adding scripts/ci/brew-ensure.sh to it
reddened app-host-execution. The addition is intended: without it the
helper is absent from .e2e-workflow and the action falls back to the
tested revision's copy.

Also assert every entry exists. A sparse-checkout of a path that is not
in the repo is silent, so a typo there would leave the action reading a
missing file and silently taking the older copy instead.

Verified: python3 tests/test_ci_e2e_compilation_cache.py, 25 tests OK.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…#15887)

* test: cover agent activity after turn completion

* fix: keep agent status running while hooks show activity

* fix: ignore ambiguous late tool activity

* Add subagent and waiting agent states to the sidebar

The compact status glyph had one "running" state, so a pane running a
fan-out of subagents, a pane parked on a background command, and a pane
typing a reply all looked identical. Two of those are worth telling
apart: subagent work is the loudest thing an agent does, and a pane
waiting on a deterministic wakeup is not asking for anything.

Claude's hooks now report what a running pane is running on through a
new `set_status --work=running|subagents|waiting` option:

- PreToolUse with `tool_name` of `Task` reports subagents. A Task call
  blocks the parent inside the tool until its subagents finish, so the
  state holds for exactly that span and the next parent hook clears it.
  No counter to drift.
- Stop with a live background task or scheduled wakeup reports waiting
  instead of running. A re-entrant Stop stays running: that is the agent
  itself still going.

The work state rides alongside the agent lifecycle rather than inside
it. A waiting pane keeps reporting a running lifecycle on purpose, so
hibernation can never SIGTERM live background work; the work state is
presentational only, and the resolver reads it before the lifecycle
branch. Waiting wins only when every agent in the workspace reports it,
so one agent still working keeps the row running.

Glyphs: subagents is a pulsing gray connected-points symbol, waiting is
a still gray hourglass. Waiting does not pulse, because the agent is
parked and a pulsing hourglass would claim otherwise. Both are
configurable through `sidebar.compactStatusIcons`, and both reach the
non-compact rows through the icon the hook sends.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
(cherry picked from commit b4bee23)

* Add dogfood tours for the new sidebar agent work states

Two tours over the same five workspaces (subagents, waiting, running,
needs input, idle): one with the compact glyph on, one with it off so
the metadata rows show the icons the hooks send.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 5904e8a)

* Fix the work-state compile breaks and the shared-key waiting regression

CI on 5904e8a caught two compile breaks the branch shipped with: the two
`shouldReplaceStatusEntry` call sites in SidebarOrderingTests never gained
the new `workState` argument, and a new control-socket test called
`hasPrefix` on an optional response. `everyIconSlotHasADistinctState` also
still pinned 11 icon slots against the 13 the branch now has. The cmuxTests
target could not build, so none of the branch's own tests ran.

The review that ran alongside it found three behavioral defects:

Subagents never appeared on a current Claude Code. The PreToolUse row
matched only `tool_name == "Task"`, and 2.x sends `Agent` for the same
spawn. Both names now count, the way `AgentChatSessionRegistry.isTaskSpawn`
already handles it for the mobile child-run tracker.

An hourglass could cover a pane that was still working. Status entries are
keyed per workspace while lifecycle states are keyed per panel, so two
Claude panes in one workspace share one `claude_code` entry and the second
to report wins. Waiting now also requires that every running lifecycle is
covered by a waiting report, so a sibling pane mid-tool-call keeps the row
running. Two panes both waiting under one key read as running, which is
the conservative direction.

The work state is now listed by `list_status` and `sidebar_state` as
`work=<state>`, so the state behind the glyph is observable instead of
screenshot-only.

Also: the doc comment promised that an unknown work state degrades to a
plain running row, while the socket rejects the whole `set_status` the way
it already rejects an unknown `--format`; the comment now describes what
the code does. `SidebarAgentWorkState.parse` dropped a `_`/`-` pass that
no input could reach and a singular `subagent` alias the socket rejects,
so the two parses accept the same set. The glyph header and
docs/configuration.md listed Running above Waiting while the resolver
checks Waiting first.

Tests: the renamed spawn tool, the two-pane shared-key case both ways, two
agents both parked, the listing line, and a pin on the raw values the
sidebar and control-socket copies of the wire contract share.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 1bad27d)

* Append the work option last so it does not split an older command prefix

CI caught this on the app-host lane:
CLINotifyProcessIntegrationRegressionTests.testClaudePromptSubmitFrom
NewSessionCanReplaceStoppedSession asserts the prompt-submit command as a
prefix through `--tab=`, and `--work=running` was being inserted between
`--color=` and `--tab=`, so the prefix no longer matched. Three assertions
in tests/test_claude_hook_clear_running_status.py use the same contiguous
fragment and would have failed on their own lane for the same reason.

None of those four assertions is about work states; they check that
prompt-submit sets Claude running on the right tab. Options are
order-independent on the wire, since the coordinator reads a parsed option
dictionary, so the new optional one goes at the end of the command instead
and the older assertions stay intact. Updating them to expect
`--work=running` would have coupled four unrelated checks to this feature
and broken them again the next time the work state for prompt-submit
changed.

Pinned by a new test in ClaudeHookWorkStateTests: the running command must
still start with the historical prefix and must end with the work option.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(sidebar): list the work option in the socket help output

The `help` text for `set_status` was the one place that still omitted
`--work`, while the usage and error strings in both coordinator copies
already list it.

Pin the work-state ordering test through the workspace id, so it stands
in byte for byte for the prefix the older suites assert, and say in the
comment why order independence holds: every option here is `--key=value`,
which a future bare flag would not be.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep the sidebar pill on Running for a re-entrant Claude Stop

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: replay deterministic wait tools in sidebar hooks

* fix: show deterministic Claude waits in sidebar

* fix: reopen identityless fresh tool activity

* fix: gate identityless activity by hook phase

* test: model fresh identityless tool resolution

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…#15409)

* Add the window recording request and frame model to CmuxFoundation

A clip needs its parameters settled before any capture starts: the format
and its default frame rate, the scale and width caps, the crop rectangle in
window points, the file name, and how the frame size follows a window that
is resized mid-clip. None of that needs a window or a screen, so it lives in
CmuxFoundation with tests instead of inside the capture session.

WindowRecordingFrameGeometry aspect-fits each captured image into the frame
size chosen at the start, so a resize letterboxes rather than stretches, and
a crop is adopted once and then kept for the rest of the clip.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add cmux record for capturing a cmux window to mp4 or gif

An agent working in its own cmux can describe what it changed, but it cannot
show it. `cmux record start` films a cmux window, or a region of one, and
writes an mp4 or a gif that can go straight into a pull request. `cmux record
note` drops a caption into the clip as it runs, so a reader can follow what
was being done.

Capture uses ScreenCaptureKit's own-process content, so only cmux's own
windows are reachable and no Screen Recording permission is asked for.
Encoding is AVAssetWriter for mp4 and ImageIO for gif; frame times come from
when each frame was sampled, so playback matches what happened. A recording
stops itself at --max-seconds, and one runs at a time, so an agent that goes
away cannot leave a capture running.

The window.record.* socket methods run on the socket worker and are not
main-thread callable: the window being filmed has to keep drawing while the
sampler runs. They are deliberately absent from the remote relay allowlist,
which stays default-deny, because a clip is local screen content rather than
an object of the peer's workspace.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add failing tests for the window recorder's crash and early-stop paths

A review of the recorder found seven defects the existing tests missed.
These are the reds, before the fixes:

- `--fps nan`, `--fps inf` and `--fps 1e30` reach `Int(Double)` in the
  request decoder and trap, which takes the whole app down with the
  socket. The package suite now crashes with "Double value cannot be
  converted to Int" instead of finishing.
- a gif whose recording stopped long before its frame budget cannot be
  finalized, because the budget is handed to ImageIO as the number of
  images the file will contain, so the documented `--gif` flow leaves no
  file at all.
- `record stop` after a clip reached its own `--max-seconds` limit
  reported "no recording is running" rather than the finished clip.
- an unknown recording id, and a note for a clip that already stopped,
  had no coverage at all.
- the mp4 writer's no-frames path was not checked for leaving a stub
  file behind, and the two-frames-one-tick test only asserted a nonzero
  duration rather than the timestamps it exists to pin down.
- `fit` upscales a window that shrank mid-clip, which its own name says
  it does not do.

`remember` on the registry stops being private so a test can set up a
finished clip without a window on screen.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix the window recorder's crash, race and early-stop paths

Seven defects a review found, in the order they bite a caller:

- `--fps nan` and `--fps 1e30` trapped in `Int(Double)` and took the app
  down. The decoder now refuses a value that is not finite or does not
  fit an `Int`.
- a gif was created with the recording's frame budget as its declared
  image count, so a clip stopped before that budget could not be
  finalized and left no file at all. The count is a floor rather than a
  cap, and how many frames a clip ends with is only known when it stops,
  so declare one and add as many as were captured.
- `stop` could close the file while an append was suspended inside the
  session actor, which is an uncatchable AVFoundation exception for the
  mp4 writer and concurrent state for the gif writer. Appends and the
  close now take turns.
- two `record start` calls could both pass the one-at-a-time check,
  because the check was followed by the await that opens the capture.
  The slot is claimed before that await.
- `record stop` after a clip reached its own `--max-seconds` limit said
  no recording was running; it reports the finished clip, as `status`
  already did.
- the mp4 writer cancelled a writer it had never started when a clip
  captured no frames, which raises rather than returns an error.
- the output file was deleted before the first frame was written, so a
  failed recording destroyed whatever was at `--out`. Frames go to a
  hidden sibling file and only move into place once the clip closes, a
  directory or device at that path is refused, and a partial file is
  cleaned up on every failure path.

Also: a window closed mid-clip reported an empty rectangle and was
captured as a one-pixel frame for the rest of the clip, so an empty
rectangle ends the recording and keeps what came before; `fit` no longer
magnifies a window that shrank mid-clip, which is what its name always
claimed; and `window.record.*` is listed in `system.capabilities` for
local callers, where the relay policy still filters it out.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Refuse a record flag that swallowed the next flag as its value

`cmux record start --label --gif` named the clip "--gif" and recorded an
mp4, because the shared option parser takes whatever follows a flag. The
record command now hands a value that looks like a flag back to the
unexpected-arguments check, which names both, and the positional
recording id for `stop`/`status` no longer accepts one either.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Map an unusable recording output path to invalid_params

`outputNotAFile` was added to `WindowRecordingSessionError` in the fix
commit for the output-path handling, but the router's switch over that
enum was never extended, so the app target stopped compiling:

    Sources/TerminalController+WindowRecording.swift:175:13:
      error: switch must be exhaustive

Nothing on the pull request caught it. The checks that run on a pull
request build the packages and the tooling, not the app target, so this
first appeared in a dispatched UI test run:

    https://github.com/manaflow-ai/cmux/actions/runs/36412249068

The missing case is `invalid_params`: `--out` naming a directory, a
device or anything else the recorder may not replace is the caller's
parameter, not a cmux failure, and an agent that reads `internal_error`
retries instead of fixing its flag.

The mapping had no test at all, which is why a missing case could sit
here. `recordingErrorCode(for:)` is now internal and every case of the
three error types it understands is pinned, including the unrecognized
fallback. A compile error cannot be committed as a red test first, since
the test would not build either; the build log above is the red
evidence, and the test is what keeps the codes from drifting.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Share the recorder's parameter decoding with a screenshot request

`window.screenshot` takes the same region, the same scale and width caps and
the same rule about an output path as `window.record.start`, and a caller who
finds a rectangle in a clip should be able to shoot it with the same four
numbers. Written twice, the two would drift: one would round a region's edges
differently, or accept a relative path, or trap on `--max-width 1e30` where the
other refuses it.

So the decoding moves into `WindowCaptureValueDecoding`, which both requests
translate into their own public failures. The recorder's API and every message
it produces are unchanged; `WindowRecordingRequest.make` now decodes the format
first and maps the shared failure back into its own, which is why the format
still names itself in a path error.

`WindowRecordingFrameGeometry.plan` gains an overload taking a region, a scale
and a width quantum rather than a recording request, so a still can plan its
crop with the code a clip uses. H.264's even-width rounding becomes that
quantum instead of being read off the format.

`WindowRecordingLabel` takes a fallback, so a label that sanitizes away to
nothing is filed as a screenshot rather than under the word "recording", and
`WindowRecordingOutputNaming` gains the screenshot directory the DEBUG command
has always used plus a filename overload taking an extension.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Discard a timed-out record start and keep standalone -- in notes

A record start that timed out on the socket kept running and could
register a recording the caller had been told failed. The registry now
tracks each start by token; on timeout the socket handler abandons it
and waits for the session to release its writer and partial file before
answering. A start that is stopped or abandoned mid-capture no longer
opens a writer, and fail() releases a writer even after the session
reached a terminal state.

record note now strips only a leading -- terminator, and cli.usage.record
carries every catalog locale.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Scope recording limits and output naming onto WindowRecordingRequest

The package conventions lint rejects all-static public types, so the
limits and the default output naming become extensions on the request.
The mp4 test names its encoded clip length plainly so the determinism
check does not read it as a measured wall-clock time.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add `cmux shot` for screenshotting a cmux window

An agent that films a window with `cmux record` still has no way to grab
one frame of it. A clip is the wrong artifact for "what does this sheet
look like now": it is slower to produce, heavier to attach, and a
reviewer cannot read a label off it.

`window.screenshot` captures one of cmux's own windows to a png or a
jpeg, through the same ScreenCaptureKit path the recorder uses. No
Screen Recording permission is involved, so it works on a release build
and inside CI, where the DEBUG-only `screenshot` command does not exist
at all.

The two commands share every decision they could disagree about:
parameter decoding, crop and scale planning, frame capture, output-file
handling and window resolution. `cmux record --region 0,0,420,900` and
`cmux shot --region 0,0,420,900` frame the same rectangle of the same
window, so a detail spotted in a clip can be shot with the same four
numbers.

The image is encoded beside the output path and moved into place, so an
existing file there is replaced only once there is a complete image to
replace it with, and `--out` pointed at a directory comes back as
`invalid_params` rather than deleting it.

Local socket only: `window.screenshot` is not on the `cmux ssh` relay
allowlist, since it reads pixels of windows the remote session does not
own.

## Changelog

Added: `cmux shot` captures a cmux window, or a region of one, to a png
or a jpeg.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Compare recorded pixels through an Equatable value

Optional tuples are not Equatable, so the caption test did not compile.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add a cmux-capture skill for screenshots and clips

`cmux shot` and `cmux record` exist, but an agent looking for a way to show
a reviewer what a change looks like has to find them first. The skill says
which of the two fits a still and which fits a change over time, how to reuse
one `--region` rectangle across both, and what an agent recording its own
cmux needs to know: the recording stops itself, one runs at a time, captions
are drawn in as it goes, and nothing but cmux is in the frames.

`references/commands.md` carries what `--help` cannot: the response fields,
the error codes and the limits the app enforces.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add a capture topic to cmux docs

`cmux docs` is where an agent with no repository checkout looks for a topic,
so the capture skill and its command reference are reachable there too, and
from the agents topic next to the hook integrations.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Guard the capture skill against drift from the code

A doc that names a response field or a state the app no longer emits sends
an agent after something that is not there. Two such errors were in the draft:
it called the recording states `recording` or `stopped` when they are
`recording`, `finished` and `failed`, and called the screenshot `format` field
`png` or `jpg` when the response carries `jpeg` and only the file is named
`.jpg`. Both came from reading the CLI and not the payload.

The guard reads both sides: documented flags against the CLI help, documented
states and formats against their enums, documented response fields against the
keys the payload actually sets, and documented error codes against the ones
the socket methods return. It needs no build, and it runs in the skill
contract workflow when the skill or the code behind it changes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Probe the capture docs topic from the CLI contract

`test_cli_contract_help.py` runs every probe line against the built CLI, so
the new topic is checked the way the dock topic is.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(skills): complete capture discovery contract

* fix(skills): advertise capture docs in CLI help

* Fix window capture edge cases

Validate and safely clamp region geometry, including origin-anchored crops and finite values beyond integer range. Build GIFs only after their exact frame count is known, staging compressed frames on disk. Keep screenshot output promotion on the waiting socket worker so a timed-out capture cannot later overwrite the requested path.

* fix(capture): preserve directories at promotion

* fix(capture): bound frame work and caption history

* fix(capture): isolate uncooperative frame requests

* fix(capture): bound GIF staging and sampling

* Pin the capture doc claims a reviewer found wrong

Nine doc findings from a review of this skill, and three guards so the same
class of drift is caught next time rather than read past.

The corrections: the error table said `invalid_params` covered an unknown record
subcommand, which never reaches the socket; a region under 8 or over 100000
points, a relative `--out` and an `--out` whose extension contradicts the format
were all undocumented; `not_found` covers an unknown recording and a window that
was never open, not only one that closed mid-capture; `conflict` also answers a
`stop` that raced its start; `fps_effective` needs two frames spanning a nonzero
time, not just two frames.

The guards: the documented region bounds are compared against the `Double`
constants both requests enforce, every relative link in the two pages has to
resolve to a file and to a heading that exists, and every Swift file these
checks read has to be in the workflow's trigger paths. The last one was missing
`WindowRecordingRequest.swift`, so a change to the recorder's own bounds could
not have failed this test.

The link guard immediately caught a forward reference: the skill pointed at
`dogfood-scenarios.md#record-a-clip`, a heading that arrives with the dogfood
record step, so the anchor is dropped until that lands.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Wait on deadlines and make the sample schedule a value

Two CI guards refused the capture branch, and both were pointing at
something worth changing rather than at the gate being strict.

`WindowScreenshotWriteTests` waited by yielding a fixed number of times.
N yields is however long N reschedules take, so the wait tightens exactly
when the runner is busy, which is when a passing test turns into a flake.
It now yields until a five second deadline instead.

`WindowRecordingSampleSchedule` was an enum of static functions that took
the previous target back in as a parameter, so the caller held the state
and the type held none. It is now a struct the sampler owns and steps
forward, which is also why its tests can assert on `targetUptime` after a
served slot rather than only on a returned number.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Pin the writer's error codes, which nothing tests any more

`record start` opens the output file before it answers, so an `--out` the
recorder cannot create fails inside the writer rather than in the request
decoder. `TerminalController.recordingErrorCode(for:)` has no branch for
`WindowRecordingWriterError`, so every one of them falls through to
`internal_error`, and an agent told by the docs to fix its flags on
`invalid_params` reads its own unwritable path as a cmux bug and retries.

This test covered it until the branch was repaired; the case that made it
redundant, `stoppedWhileStarting`, is gone, but the writer errors are not.
It is restored first, failing, and the mapping follows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Tell a caller when the recorder could not open their clip

`WindowRecordingWriterError.setup` is the file the caller named refusing to
be created, so it answers `invalid_params`; a frame or a close that fails is
the encoder, so those stay `internal_error`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Pin the limits and codes an agent reads, before fixing them

Three claims the capture docs make that nothing checked.

`invalid_params` tells an agent to edit its flags, so it cannot also mean
"this window has nothing in it": no flag makes an unrendered window
capturable, and the mapper sent all three geometry failures down that
branch. The test now pins the split.

The sample schedule counted one slot past every elapsed slot, so a
sampler arriving exactly on a later slot threw that frame away and waited
a whole interval more for the next one. The new case asks for the slot it
is already in time for.

The limits table was checked for flag names and never for numbers, which
is how it came to promise a gif `--max-width` of 4096 against a ceiling of
1280, and to omit the two gif budgets that have no flag of their own. The
guard now reads the constants and asserts the page states each one, and
the region check compares values instead of mirroring a source line.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep an on-time frame, and stop blaming flags for an empty window

`recordingErrorCode` now answers `not_found` for `emptyWindow`, beside
the two geometry failures a parameter can fix. An agent that reads
`invalid_params` changes its flags and tries again, which no flag would
have helped here.

The schedule rounds up to the first slot at or after now instead of
adding one to the elapsed slots, so a target that is exactly `now` is
served rather than skipped, and it steps once more in the one case where
the division rounds down in binary.

Two comments claimed a window resized mid-clip never ends a recording.
That holds for the frame size, which is fixed when the clip opens, but
not for a `--region` the window has shrunk out of: the crop is planned
again per frame, planning throws, and the clip ends in state `failed`
with the frames it had. Both comments now say so.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Document the gif ceilings, and offer an index in --help

The limits table promised a gif `--max-width` of 4096 while a gif stops
at 1280, and said nothing about the two budgets a gif has and an mp4 does
not: 960 frames, and 4000000 pixels a frame. The reference page owns the
limits and `--help` owns the flag list, so the ceilings go here rather
than into a wider flag column.

Also corrected on that page: a region a shrinking window leaves behind
ends a clip, which the region section had not admitted; `conflict` no
longer covers a stop that raced its start, a case that no longer exists;
an `--out` under a read-only volume is `invalid_params`; a window with no
capturable content is `not_found`; and `rpc window.screenshot` takes a
window id, because the CLI is what resolves a ref or an index. The skill
page said `record status` reports the effective frame rate, which only
`--json` carries.

`--help` for both commands listed `--window <id|ref>` while the usage
line and the reference page both offer an index, and the CLI has always
accepted one. Same token in all nine locales, since it is a spec rather
than prose.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(recording): stop asserting how many samples the encoder writes

`twoFramesInTheSameTickStillGetIncreasingTimes` expected a two frame clip
to read back as exactly two samples, and the CI VMs write six: with no
hardware scaler available, the encoder decides how many samples a clip of
two frames a 1/600 second apart ends up with.

The property under test survives. A dropped frame, which is what an equal
presentation time used to cause, leaves a single sample, so a floor of two
still catches it, and the times still have to start at zero, increase, and
put the second sample a tick rather than a frame interval after the first.
The expectations now print the times they read, so the next surprise from
the encoder arrives with its evidence.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(recording): pin that a gif keeps the frames staged before its limit

A gif that reaches maximumStagedBytes mid-clip currently loses every frame:
finish() retries the frame it could not stage, throws again, and never creates
the destination, so the session deletes the working file. That contradicts
fail(_:), which promises "Keep whatever was captured before the failure".

Red before the fix: the finalized gif does not exist.

## Changelog

none

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(recording): pin what a mid-clip crop replan may and may not fail on

A clip's encoded frame size is fixed by its first frame, so replanning a later
frame with the gif pixel ceiling still applied ends a recording over a size
nothing encodes: a window that grows mid-clip fails with gifFrameTooLarge even
though the frame is drawn into the opening size. A region the window shrank out
of is the one resize that still has to end the clip.

Red before the fix: planCrop does not exist.

## Changelog

none

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(recording): keep a gif's staged frames when the last one cannot stage

`finish()` staged the frame it was holding without clearing `pending`, so a
clip that had already reached its staging limit threw again from the one call
that is supposed to finalize it. The session's `fail(_:)` promises the caller
"whatever was captured before the failure", and that promise was being broken
for exactly the recordings that need it.

Drop the one frame that has no room and finalize the frames before it. A clip
with nothing staged still throws, so a recording with no frames still reports
why it has none.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(recording): re-crop a window that grows mid-clip instead of ending the gif

The per-frame replan called the first-frame planner, which applies the gif
pixel ceiling. A window that grew past that ceiling mid-clip therefore ended
the recording over a frame size nothing encodes: the encoded size was fixed by
the first frame, and the grown crop was only ever going to be drawn into it.

`planCrop` plans the crop with no pixel ceiling, and the session uses it for
later frames. The one resize that still ends a clip is a `region` the window
has shrunk out of, because no rectangle is left to sample.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(cli): give the `screenshot` alias its contract row

`check-cli-contract-verbs` reads dispatched verb strings, and the shot arm
dispatches both `shot` and `screenshot`. Only `shot` had a row, so the guard
failed the whole run with "top-level verb screenshot is in no command table".

The first cell may name a verb and its aliases, so both go in the one row.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: validate screenshot requests and capture docs

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test: reproduce answered agent notification ring

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: clear answered agent notification rings

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: clear Codex answered notification rings

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: fence Codex prompt notification retirement

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: normalize agent session fallback matching

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: stamp ordered Codex progress hooks

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: limit Codex progress stamps to safe hooks

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: clear agent rings on terminal input

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: clear agent rings in dock terminals

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: remove unused semantic notification binding

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: isolate agent notification mutation fixtures

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: clear queued mutations between notification fixtures

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: fence superseded agent prompt retirement

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: preserve newer same-session prompts

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: index unread agent prompts

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…anaflow-ai#15989)

* test: reproduce late Agent Chat startup stealing session focus

* fix: keep late Agent Chat startup replies from changing selection

* Make Agent Chat fork and handoff actions recoverable (manaflow-ai#15994)

* test: reproduce stuck Agent Chat fork and handoff lifecycle

* fix: make Agent Chat fork and handoff actions recoverable

* fix(agent-chat): deduplicate reconnecting session actions

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(agent-chat): keep main's pinned tsc in the check script

The main merge resolved agent-chat/package.json toward this branch's
`bun x tsc --noEmit`, but main had since pinned typescript 5.9.2 as a
devDependency and added tests/test_ci_guard_workflow_structure.py, which
asserts the exact check script. CI fast guards failed on that assertion.

The pinned dependency came through the merge, so ./node_modules/.bin/tsc
exists; only the script string was stale.

Verified: python3 tests/test_ci_guard_workflow_structure.py passes, and
`bun run check` in agent-chat is 106 pass / 0 fail across 18 files.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…w-ai#14852)

* Checkpoint terminal scrollback so a crash no longer loses it

The 8 s session autosave never captures scrollback, so after a crash or
SIGKILL every terminal restored empty; only clean quit, power-off and
update relaunch persisted it (manaflow-ai#2016, manaflow-ai#2194).

Add slow, bounded scrollback checkpoints next to the primary snapshot:
- at most every 60 s, only after 5 s without typing, driven from the
  existing autosave timer;
- only terminals that produced PTY output since their last capture,
  flagged by one relaxed atomic load per PTY read in the existing tee;
- at most 3 Ghostty VT exports per checkpoint, one per main-queue turn,
  stopping early on typing or after 50 ms of main-thread capture time;
- truncation, encoding and writes on a utility queue, one file per
  terminal under session-<bundle>-scrollback/, pruned to live panels;
- the same eligibility gates as the quit path (running command,
  hibernated agent), with stale checkpoints deleted.

After an unclean exit, startup restore fills each terminal's missing
scrollback from its checkpoint; the newer of snapshot and checkpoint
wins, and the restore path still applies its own replay gates.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Bound checkpoint capture cost and keep restored scrollback across crashes

Review follow-up for the scrollback checkpoints:

- Split the capture: only Ghostty's VT export runs on main; reading the
  export file, CRLF normalization, the 4000-line tail, truncation,
  encoding and the write run on the utility queue. A terminal whose
  export alone exceeded the 50 ms budget, or failed, is skipped for
  10 minutes.
- Seed checkpoints from restored scrollback when a restore completes,
  keyed by the restored panel ids and without a VT export, so a second
  crash before the next checkpoint keeps it.
- Record scrollbackCapturedAt on snapshots saved with scrollback. A
  newer scrollback-bearing save wins over a checkpoint even when it
  deliberately omitted a terminal's scrollback; the 8 s autosave (no
  marker) still yields to checkpoints.
- Re-mark a terminal pending when its export read or file write fails.
- Keep a terminal pending for one more checkpoint when output arrived
  between planning and capture, since the PTY tee runs before Ghostty
  parses those bytes.
- Disable checkpoints under automated test runs, like session restore.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Discard checkpoint exports on quit and seed only after a crash

- A checkpoint interrupted by quit or restore now deletes the export
  files it already wrote synchronously on main (one unlink each) and
  leaves those terminals pending, instead of handing them to the
  utility queue, which may not run before exit.
- Seed checkpoints from restored scrollback only after an unclean
  previous launch, not on clean launches or manual reopen, to avoid
  rewriting up to 400 KB per terminal needlessly.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix overlapping access in the scrollback checkpoint merge

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Discard checkpoints after a clean exit so a later crash cannot restore them

A clean quit saves scrollback into the snapshot but left the checkpoint
files behind until the next launch's first checkpoint pruned or rewrote
them. Restore reuses snapshot panel ids (workspace and dock terminals),
so if that next launch crashed first, its 8 s autosave (no scrollback,
no capture marker) matched the old records and the crash restore
replayed the earlier launch's scrollback over what the quit saved.

Startup now deletes the checkpoint directory when the previous launch
exited cleanly, before the restore decision so a skipped restore also
drops them. After an unclean exit the checkpoints are that launch's own
and are still merged and kept until the restored terminals are seeded.

Moves the checkpoint coordinator setup into the AppDelegate checkpoint
extension to keep AppDelegate.swift's growth down.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Inject one scrollback checkpoint activity owner and delete removals before returning

The PTY tee bridge, the checkpoint coordinator and the persist step now share
one `TerminalScrollbackCheckpointActivity` owned by the composition root
(`GhosttyApp.terminalScrollbackCheckpointActivity`) instead of a `.shared`
singleton, so a caller cannot hand the coordinator a different instance than
the one the tee writes.

Planned removals (a panel that stopped being eligible) are now applied with
`queue.sync` after earlier queued writes, and only captures and pruning go
async. A crash before the async block ran used to leave the old checkpoint on
disk for the next unclean restore to merge.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test scrollback checkpoint output interleaving

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Make scrollback checkpoints generation-safe

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Preserve scrollback for idle agent sessions

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Match checkpoint agent liveness policy

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ai#15928)

* test: cover Ctrl-letter key event encoding

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: encode synthetic Ctrl-letter keys

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: compare named Ctrl-letter bytes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: report named key releases

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: avoid shadowing Ghostty key event type

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: retain synthetic key press capture

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: match Ghostty legacy Ctrl-I and Ctrl-M

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: restore bonsplit pin for macOS compile

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
manaflow-ai#15420)

* Move browser and tmux OSC filter tests into their package test targets

BrowserSystemProxyMirror, BrowserUserAgentPolicy, BrowserNavigationDecisionHandler,
BrowserHiddenWebViewDiscardManager and BrowserHTTPBasicAuthPromptCoordinator live in
CmuxBrowser, and RemoteTmuxNotificationOSCFilter in CmuxRemoteSession (manaflow-ai#13787), but
their 68 tests compiled into cmuxTests and ran inside the launched app host. They use
only package API, so they move unchanged except for dropping the cmux import.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Drop the blank line the removed import block left

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep app shard metadata aligned with package tests

* test: give split fixtures deterministic geometry

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add failing tests for agent wake verification

The tests need new types, so this commit also adds minimal stubs:
AgentWakeVerificationState never changes state, and the Workspace entry
points (beginAgentWakeVerification, failAgentWakeVerification,
dismissAgentWakeFailure) do nothing. The state machine and Workspace
tests fail against these stubs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Check that a woken agent came back

Waking a hibernated agent types its resume command into a fresh shell,
but nothing checked whether the agent started. A missing launcher or a
gone session left the pane at a shell prompt with no sign of failure.

A wake now starts a check for the pane. An agent hook reporting a PID or
a lifecycle state counts as success. The check fails when the resume
command returns to the prompt first, or when nothing reports within 90
seconds and no live agent process is found. A failure shows a banner on
the pane (Retry, Show command with Copy, Close), an "Agent didn't
resume" sidebar row, and one notification feed entry. The row is kept
out of the saved session. Closing the pane, hibernating it again, or a
later agent report clears the check.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Only fail quick resume exits and never retype into a running program

A resume command that returns after 20 seconds ran the agent, which the
user then quit; for agents without hooks that was reported as a failed
wake. It now counts as resumed. Retry is hidden, and refused, while a
command still runs in the pane, since the text would go to that program.
The failed-wake row survives a sidebar reset, and a retired workspace
skips the deadline.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Confirm wakes only from the woken agent's own evidence

Hook reports now count only when they come from the woken agent and
name their pane explicitly; the focused-pane fallback no longer confirms
a wake. A resume command that ran a while is no longer taken as proof:
a pending check samples the pane for a live agent process every few
seconds instead, and a command that ends before any confirmation fails
the wake. The verification deadline constants are nonisolated, which
fixes the new Swift warning.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
teamleaderleo and others added 2 commits September 30, 2026 12:15
* test(cloud): cover the reads that find a Cloud machine gone

A status read and an access preflight both move the row to `destroyed`
themselves when the provider no longer has the machine. Once either does,
`destroyVm` can never see the row again (its lookup skips destroyed rows)
and the provider-status cron skips it too, so the model-plane revoke and
the `vm.destroyed` usage event that the cron performs for the same
transition have to happen at these write sites as well.

Fails today: both retire the row and record nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): finish the destroy work where a machine is first found gone

Three places retire a Cloud machine's row after the provider says it no
longer has the machine: the status read, an access operation's resume
preflight, and the stats read. Each wrote `status = destroyed` and stopped
there, while the reconcile cron doing the same transition also revoked the
machine's model-plane tokens and recorded a `vm.destroyed` ledger event.

That difference was permanent, not a race to lose. `destroyed` is terminal:
`findUserVm` hides such a row from every destroy request and
`reconciliationCandidates` drops it from the cron, so whichever of the three
got there first left a machine that never appears as destroyed in the ledger
and never has its route tokens marked revoked, with nothing able to finish
the job afterwards.

All three now go through one `applyObservedProviderStatus` that performs the
write and, when the write lands on `destroyed`, the same revoke and ledger
event as the cron. The status route hands `getVm` the model-plane revoker it
already builds for delete. The new `provider_status_*` destroy reasons join
the analytics allowlist so the event says which read noticed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): derive the observed status inside the shared destroy path

Review found that the previous commit took the row's new status from each
caller. That let the access preflight and the stats read hardcode
"destroyed" for a provider 404, including for a machine with a persistent
home volume, which observedDbStatus maps to "paused" because the compute
is gone and the machine is not. Those rows were terminalized and billed a
vm.destroyed that never happened, and nothing revisits a terminal row to
take either back.

The status is now derived inside applyObservedProviderStatus, so every
entrypoint agrees about what a 404 means. reopenBaseIfProviderDeleted is
the one caller that must override it, and passes forceStatus with the
reason: its row is a Base's active generation, and leaving it paused would
hand the same dead provider id back on every later open.

Also threads the model-plane revoker through getVmStats from its route,
covers the stats entrypoint including the home-volume case, and fixes the
two test-file regressions the previous commit shipped.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* vm: say what a paused 404 row means for credentials

The access preflight carried a comment claiming a revoke was pointless
there because the credentials were already inert. That is false for the
case this branch introduces: authenticateRouteToken and
authenticateVmAuthorization both accept `paused`, so route tokens on a
volume-backed machine stay valid after this write, for the rest of their
30-day lifetime.

Keep the behavior, which is what getVm and the reconcile cron already do
for the same observation, and replace the comment with what is true. Also
drop the stale "next fleet refresh drops it" line, correct "seven call
sites" to eight, stop asserting a resurrection path that is not
implemented, and type usageEventSource as VmDestroySource so an unknown
source cannot silently degrade in PostHog.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(cloud): reproduce lost destroy ledger retry

* fix(cloud): retire missing provider compute durably

* fix(cloud): durably drain observed destroy cleanup

* fix(cloud): keep destroy cleanup queue fair and indexed

* fix(vm): make observed cleanup retries index-safe

* test(cloud): reproduce account deletion cleanup loss

* fix(cloud): retain pending destroy cleanup on account deletion

* fix(cloud): transfer account cleanup to durable outbox

* fix(cloud): reject malformed cleanup outbox rows

* fix(cloud): validate cleanup outbox shape

* Persist explicit destroy cleanup for retry

* fix cloud observed cleanup retries and concurrent migration

* fix migration retry for invalid concurrent indexes

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@cursor

cursor Bot commented Sep 30, 2026

Copy link
Copy Markdown

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

teamleaderleo and others added 2 commits September 30, 2026 19:16
* feat: surface configured actions and launchers

* fix(actions): preserve surface button action identity

* fix(actions): report exact placements and global edit target

* test(actions): pin exact surface placement identity

* test(config): preserve surface action references

* fix(actions): drop hidden surface action references

* test(config): drop references for hidden surface buttons

* fix: localize configured actions empty state

* fix: sync upstream CmuxFoundation import repair (manaflow-ai#13236)

Applies upstream 968cef2 to unblock this PR's app-host compilation.

* fix(actions): localize discovery metadata labels

* fix: localize actions discovery titles and open button

Give the Actions menu item and dialog their own whole-string keys instead of
borrowing the titlebar-layout debug label and concatenating, and label the
open button with the existing Open cmux.json key rather than a raw path.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: document French identity for the actions discovery titles

"Actions" is spelled the same in French, and cmux.json is a file name.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

needs a call Finished and held for a team design or product decision (see #13742 and the gallery in #15427)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants