Skip to content

feat: auto-name tabs from Claude Code conversation content - #2041

Closed
Horacehxw wants to merge 4 commits into
manaflow-ai:mainfrom
Horacehxw:feat/auto-tab-naming-from-transcript
Closed

Horacehxw wants to merge 4 commits into
manaflow-ai:mainfrom
Horacehxw:feat/auto-tab-naming-from-transcript

Conversation

@Horacehxw

@Horacehxw Horacehxw commented Mar 24, 2026 •

Copy link
Copy Markdown

Summary

  • Auto-generates tab titles from Claude Code conversation content using keyword frequency extraction
  • Runs in the existing claude-hook stop handler with zero additional latency (transcript is already being read)
  • Supports multilingual conversations (CJK + Latin tokenization)
  • Respects user /rename — if a custom title exists in the transcript, auto-naming is skipped

Motivation

Currently, cmux tabs running Claude Code sessions show generic names (e.g., the shell command or directory). When working with multiple Claude Code sessions across tabs, it's hard to distinguish which tab is doing what without manually /rename-ing each one.

This change automatically extracts meaningful keywords from user messages in the Claude Code transcript and uses them as the tab title. For example, a conversation about "debugging cmux hooks" would auto-name the tab something like "cmux hooks debugging".

How it works

  1. readTranscriptSummary now also extracts the first 3 and last 3 user messages (in addition to the existing lastAssistantMessage)
  2. generateTabTitle uses regex-based tokenization to extract CJK runs and Latin words, filters stop words (English + Chinese), and picks the top keywords by frequency
  3. The title is capped at 25 characters and 2-5 words
  4. If the user has /renamed the session (detected via custom-title entries in the transcript JSONL), auto-naming is completely skipped

Design decisions

Decision Rationale
Keyword extraction vs AI summary Zero latency, no external API dependency, runs in the synchronous hook path
Stop words in EN + ZH Most cmux users are English or Chinese speakers based on repo demographics
First 3 + last 3 user messages Captures both the initial topic and recent focus shifts
25 char / 5 word limit Fits comfortably in the cmux sidebar tab width
Skip on custom-title User intent should always take precedence over auto-generated names

Test plan

  • Verify tab auto-names after first Claude Code response in a cmux terminal
  • Verify /rename custom title is preserved (auto-naming skipped)
  • Verify CJK conversation generates CJK tab title
  • Verify English conversation generates English tab title
  • Verify existing notification behavior is unchanged

Summary by cubic

Automatically names Claude Code tabs from recent conversation keywords so you can tell sessions apart. Also stabilizes UI test launch and fixes portal geometry after window restore.

  • New Features

    • Generates short tab titles from the first/last user messages using keyword frequency.
    • Runs in the existing claude-hook stop handler with no added latency; supports CJK and Latin.
    • Skips auto-naming if a /rename custom title is found in the transcript.
  • Bug Fixes

    • Resyncs terminal portal geometry after restore-time bind so queued layout shifts (sidebar/splits) don’t leave stale hit-targets.
    • Improves UI test reliability: retries app activation on launch and re-runs the display resolution UI test in CI on the known activation flake.

Written for commit 275ee5c. Summary will update on new commits.

Summary by CodeRabbit

  • New Features

    • Tabs automatically rename based on conversation content when transcripts are available and no custom title is set.
  • Tests

    • Improved CI test reliability with retry logic for UI tests.
    • Enhanced UI test window stabilization during launch.

austinywang and others added 4 commits March 22, 2026 21:04
…anaflow-ai#1973)

* test: cover queued restore-time terminal portal shift

* fix: resync terminal portal after restore-time bind
Add keyword-based tab auto-naming to the `claude-hook stop` handler.
When Claude Code finishes a turn, the transcript is analyzed to extract
the most frequent meaningful keywords from user messages, which are used
to generate a short (2-5 word, max 25 char) tab title.

Key design decisions:
- Uses keyword frequency extraction (not AI/LLM) for zero-latency naming
- Supports CJK + Latin tokenization for multilingual conversations
- Respects user /rename: if a custom-title is detected in the transcript,
  auto-naming is skipped entirely
- Extracts first 3 + last 3 user messages to capture both the initial
  topic and recent focus shifts
- Filters common stop words in English and Chinese

This runs synchronously in the existing stop hook path with negligible
overhead since the transcript is already being read for notification
summaries.
@vercel

vercel Bot commented Mar 24, 2026

Copy link
Copy Markdown

@Horacehxw is attempting to deploy a commit to the Manaflow Team on Vercel.

A member of the Team first needs to authorize it.

@coderabbitai

coderabbitai Bot commented Mar 24, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

This pull request introduces automatic tab renaming from transcript summaries, enhances UI test stability through retry logic with window/foreground activation, adds external geometry synchronization after terminal bind operations to handle layout shifts, and implements CI retry logic for a display resolution regression test with output capture.

Changes

Cohort / File(s) Summary
CI Workflow & Display Test
.github/workflows/ci.yml
Replaced direct xcodebuild test invocation with output capture via tee and PIPESTATUS-based exit code extraction. Adds conditional retry on "Failed to activate application / Running Background" errors, with process cleanup (pkill) and 2-second delay between attempts.
Transcript Auto-Titling
CLI/cmux.swift
Expanded TranscriptSummary struct to track first/last user messages and custom title presence. Implemented generateTabTitle() using keyword extraction and stop-word filtering. Added automatic tab renaming via tab.action v2 requests during transcript completion and Claude hook idle/stop events.
UI Test Launch Stabilization
Sources/AppDelegate.swift
Introduced stabilizeUITestLaunchWindowAndForeground(attempt:) with up to 20 retries (0.25s apart) for window/foreground activation. Includes fallback window creation, macOS version-aware activation, and per-retry diagnostics. Centralizes activation logic in activateUITestAppIfNeeded().
Terminal Window Geometry Sync
Sources/TerminalWindowPortal.swift
Added scheduleExternalGeometrySynchronize() call after bind() to handle post-bind ancestor layout shifts and prevent stale portal hit-testing.
Geometry Sync Test Coverage
cmuxTests/TerminalAndGhosttyTests.swift
Added testBindQueuesExternalGeometrySyncForQueuedLayoutShift test verifying that bind schedules external geometry sync when ancestor layout changes occur after binding.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~50 minutes

Possibly related PRs

Poem

🐰 Tabs rename themselves with wisdom gleaned,
Windows stabilize through retries keen,
Geometry syncs when layouts shift,
Tests capture output—CI's gift! ✨
No stale portals shall interfere,
Smart transcripts make titles appear.

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Title check ✅ Passed The title accurately summarizes the main change: automatic tab naming from Claude Code conversation content, which is the primary feature added across the codebase.
Description check ✅ Passed The description covers all required template sections: summary with what and why, testing approach with manual verification checklist, and includes a detailed design rationale table and test plan.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
⚔️ Resolve merge conflicts
  • Resolve merge conflict in branch feat/auto-tab-naming-from-transcript

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@greptile-apps

greptile-apps Bot commented Mar 24, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR adds automatic tab titling for Claude Code sessions in cmux by extracting high-frequency keywords from user messages in the conversation transcript, and separately hardens the display-resolution UI regression CI job with a one-shot retry for a known launch-activation flake, plus a session-restore geometry fix in TerminalWindowPortal.

Key changes:

  • readTranscriptSummary now collects the first 3 and last 3 user messages and detects custom-title JSONL entries (written by /rename) so auto-naming is suppressed when the user has set a custom title.
  • generateTabTitle tokenizes combined user messages using a CJK (\u4e00–\u9fff) + Latin regex, filters a hardcoded EN/ZH stop-word list, ranks by frequency, and returns the top 2–5 keywords capped at 25 characters.
  • Logic bug: firstUserMessages and lastUserMessages overlap for conversations with ≤5 user messages. Concatenating both arrays without deduplication double-counts tokens from the overlapping messages (worst case: messages 2–3 of a 4-message conversation get 2× the frequency weight of messages 1 and 4), skewing keyword ranking toward middle messages instead of the initial topic or most recent focus.
  • The deduplication block (seen/deduped) after building top is unreachable dead code — freq keys are already unique lowercase strings, so no duplicate can ever appear in top.
  • The NSRegularExpression is compiled fresh on every generateTabTitle call; promoting it to a static let would avoid repeated compilation overhead.
  • AppDelegate gains a 20-attempt stabilization loop for XCTest launch windows, scoped to #if DEBUG with a macOS 14 availability guard for the activation API.
  • TerminalWindowPortal.bind now schedules a deferred external geometry sync to handle ancestor layout shifts during session restore, with a new unit test verifying the behavior.

Confidence Score: 4/5

  • Safe to merge after fixing the message-overlap bias in generateTabTitle; the CI, portal geometry, and AppDelegate stabilization changes are clean.
  • The auto-naming feature is well-motivated and the overall design is sound. The overlap bug (double-counting messages 2–3 in a 4-message conversation) produces biased keyword rankings for the most common short-session case, which is the primary user path for this feature. The fix is a one-liner. The dead-code dedup block and per-call regex compile are style issues that don't affect correctness. The three other changed files (CI workflow, AppDelegate, TerminalWindowPortal + test) are all solid.
  • CLI/cmux.swift — specifically the generateTabTitle function and the allMessages construction at line 10930.

Important Files Changed

Filename Overview
CLI/cmux.swift Core feature implementation: extends readTranscriptSummary to collect first/last user messages and detect custom-title entries, adds generateTabTitle with regex-based CJK+Latin keyword extraction, and wires the result into the claude-hook stop handler. Contains a logic bug where firstUserMessages and lastUserMessages overlap for ≤5-message conversations, biasing keyword frequency toward middle messages. Also includes unreachable deduplication code and a per-call regex compile.
.github/workflows/ci.yml Wraps the display-resolution UI regression test in a retry function so it re-runs once when the known "Failed to activate application — Running Background" launch-activation flake is detected. Clean, targeted change with no logic issues.
Sources/AppDelegate.swift Replaces the inline afterForceWindow DispatchQueue block with stabilizeUITestLaunchWindowAndForeground, a polling retry loop (up to 20 attempts × 250 ms) that waits for a visible key window before activating the app. Uses macOS 14.0 availability guard to drop the deprecated activateIgnoringOtherApps option on newer OS versions. Well-structured and scoped to #if DEBUG.
Sources/TerminalWindowPortal.swift Adds a scheduleExternalGeometrySynchronize() call in the bind path so that session/window restore ancestor layout shifts (sidebar width, split positions) are caught after the initial bind tick. Small, targeted, and covered by the new unit test.
cmuxTests/TerminalAndGhosttyTests.swift Adds testBindQueuesExternalGeometrySyncForQueuedLayoutShift — a runtime test that binds a hosted view, shifts its ancestor container on the main queue, waits 50 ms, and asserts portal hit-testing resolves to the new position. Tests observable behavior through an executable path, consistent with the project's test quality policy.

Sequence Diagram

sequenceDiagram
    participant CC as Claude Code
    participant Hook as cmux CLI<br/>(claude-hook stop)
    participant TS as readTranscriptSummary
    participant GT as generateTabTitle
    participant Server as cmux App<br/>(tab.action rename)

    CC->>Hook: hook event: stop (transcript path)
    Hook->>TS: read JSONL transcript
    TS->>TS: collect firstUserMessages[0..2]<br/>lastUserMessages (ring buffer, max 3)<br/>detect custom-title entries
    TS-->>Hook: TranscriptSummary { firstUserMessages, lastUserMessages, hasCustomTitle }
    alt hasCustomTitle == true
        Hook-->>CC: skip auto-rename (user /rename takes precedence)
    else hasCustomTitle == false
        Hook->>GT: generateTabTitle(transcript)
        GT->>GT: combine first+last messages<br/>tokenize CJK runs + Latin words<br/>filter stop words (EN+ZH)<br/>rank by frequency → top 5<br/>build title ≤25 chars, 2–5 words
        GT-->>Hook: tabTitle (or nil if no keywords)
        Hook->>Server: tab.action rename { workspace_id, surface_id, title }
        Server-->>Hook: OK
    end
    Hook->>Server: setClaudeStatus → Idle
Loading

Reviews (1): Last reviewed commit: "feat: auto-name tabs from Claude Code co..." | Re-trigger Greptile

Comment thread CLI/cmux.swift
/// Generate a short tab title from user messages using keyword extraction.
/// Preserves the conversation language (works with CJK, Latin, etc).
private func generateTabTitle(from transcript: TranscriptSummary) -> String? {
let allMessages = transcript.firstUserMessages + transcript.lastUserMessages

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 firstUserMessages and lastUserMessages overlap for short conversations

When there are ≤5 user messages, some or all messages appear in both firstUserMessages and lastUserMessages. Concatenating them without deduplication inflates keyword frequencies for those overlapping messages:

  • 3 user messages: every message is counted twice — ranking is effectively unchanged, but work is doubled.
  • 4 user messages: firstUserMessages = [m1, m2, m3], lastUserMessages = [m2, m3, m4], so allMessages = [m1, m2, m3, m2, m3, m4]. Tokens from m2 and m3 are counted at 2× the weight of m1 and m4.
  • 5 user messages: m3 is in both windows and receives a 2× frequency boost over m1, m2, m4, m5.

This biases keyword scoring toward middle-of-conversation messages, potentially surfacing stale topics over the most recent focus (m4/m5).

The simplest fix is to deduplicate in generateTabTitle before building allMessages:

// Remove messages already captured by firstUserMessages so they're not double-counted.
let uniqueLast = transcript.lastUserMessages.filter { !transcript.firstUserMessages.contains($0) }
let allMessages = transcript.firstUserMessages + uniqueLast

(Or collect indices rather than values during parsing so deduplication is exact even when two messages happen to have identical truncated text.)

Comment thread CLI/cmux.swift
Comment on lines +10978 to +10987
// Build title: 2-5 words, max 25 chars
var seen = Set<String>()
var deduped: [String] = []
for word in top {
let low = word.lowercased()
if !seen.contains(low) {
seen.insert(low)
deduped.append(word)
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Deduplication block is unreachable dead code

freq is a [String: Int] dictionary whose keys are already lowercased strings (built with let low = token.lowercased()). top is produced by sorting that dictionary and mapping each entry back through casing[$0.key] ?? $0.key — every entry has a unique lowercase key. Therefore seen.contains(low) is never true, and deduped always ends up identical to Array(top).

The entire seen/deduped block can be replaced with a direct assignment:

Suggested change
// Build title: 2-5 words, max 25 chars
var seen = Set<String>()
var deduped: [String] = []
for word in top {
let low = word.lowercased()
if !seen.contains(low) {
seen.insert(low)
deduped.append(word)
}
}
let deduped = top

Comment thread CLI/cmux.swift

// Tokenize: CJK runs (2+ chars) or Latin word tokens
var tokens: [String] = []
let pattern = try? NSRegularExpression(pattern: "[\\u4e00-\\u9fff]{2,}|[a-zA-Z][a-zA-Z0-9]*", options: [])

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Regex compiled fresh on every generateTabTitle call

NSRegularExpression(pattern:options:) performs full pattern compilation on each invocation. Since the pattern is a constant, it should be compiled once and reused — e.g., as a static let or a private computed property at struct scope:

private static let tokenPattern: NSRegularExpression? = try? NSRegularExpression(
    pattern: "[\\u4e00-\\u9fff]{2,}|[a-zA-Z][a-zA-Z0-9]*",
    options: []
)

This is low-impact today (the hook runs only on session stop), but it is unnecessary per-call overhead.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

3 issues found across 5 files

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="Sources/AppDelegate.swift">

<violation number="1" location="Sources/AppDelegate.swift:2425">
P2: Nested UITest retry loops can overlap and enqueue redundant main-thread retries, increasing flakiness and timing churn.</violation>

<violation number="2" location="Sources/AppDelegate.swift:2448">
P2: macOS 14+ UI-test activation dropped `.activateIgnoringOtherApps`, so retrying on `!isActive` can fail to ever foreground the app.</violation>
</file>

<file name="cmuxTests/TerminalAndGhosttyTests.swift">

<violation number="1" location="cmuxTests/TerminalAndGhosttyTests.swift:2993">
P2: New test uses a fixed 50ms wait for async coordination, which can make geometry assertions flaky under CI/scheduler jitter.</violation>
</file>

Reply with feedback, questions, or to request a fix. Tag @cubic-dev-ai to re-run a review.

Comment thread Sources/AppDelegate.swift
openNewMainWindow(nil)
}

moveUITestWindowToTargetDisplayIfNeeded()

@cubic-dev-ai cubic-dev-ai Bot Mar 24, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: Nested UITest retry loops can overlap and enqueue redundant main-thread retries, increasing flakiness and timing churn.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Sources/AppDelegate.swift, line 2425:

<comment>Nested UITest retry loops can overlap and enqueue redundant main-thread retries, increasing flakiness and timing churn.</comment>

<file context>
@@ -2416,6 +2411,46 @@ final class AppDelegate: NSObject, NSApplicationDelegate, UNUserNotificationCent
+            openNewMainWindow(nil)
+        }
+
+        moveUITestWindowToTargetDisplayIfNeeded()
+        activateUITestAppIfNeeded()
+
</file context>
Fix with Cubic

Comment thread Sources/AppDelegate.swift
window.orderFrontRegardless()
}
if #available(macOS 14.0, *) {
NSRunningApplication.current.activate(options: [.activateAllWindows])

@cubic-dev-ai cubic-dev-ai Bot Mar 24, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: macOS 14+ UI-test activation dropped .activateIgnoringOtherApps, so retrying on !isActive can fail to ever foreground the app.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Sources/AppDelegate.swift, line 2448:

<comment>macOS 14+ UI-test activation dropped `.activateIgnoringOtherApps`, so retrying on `!isActive` can fail to ever foreground the app.</comment>

<file context>
@@ -2416,6 +2411,46 @@ final class AppDelegate: NSObject, NSApplicationDelegate, UNUserNotificationCent
+            window.orderFrontRegardless()
+        }
+        if #available(macOS 14.0, *) {
+            NSRunningApplication.current.activate(options: [.activateAllWindows])
+        } else {
+            NSRunningApplication.current.activate(options: [.activateAllWindows, .activateIgnoringOtherApps])
</file context>
Suggested change
NSRunningApplication.current.activate(options: [.activateAllWindows])
NSRunningApplication.current.activate(options: [.activateAllWindows, .activateIgnoringOtherApps])
Fix with Cubic

window.displayIfNeeded()
}

RunLoop.current.run(until: Date().addingTimeInterval(0.05))

@cubic-dev-ai cubic-dev-ai Bot Mar 24, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: New test uses a fixed 50ms wait for async coordination, which can make geometry assertions flaky under CI/scheduler jitter.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At cmuxTests/TerminalAndGhosttyTests.swift, line 2993:

<comment>New test uses a fixed 50ms wait for async coordination, which can make geometry assertions flaky under CI/scheduler jitter.</comment>

<file context>
@@ -2939,6 +2939,83 @@ final class TerminalWindowPortalLifecycleTests: XCTestCase {
+            window.displayIfNeeded()
+        }
+
+        RunLoop.current.run(until: Date().addingTimeInterval(0.05))
+
+        let shiftedAnchorFrameInWindow = anchor.convert(anchor.bounds, to: nil)
</file context>
Fix with Cubic

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

🧹 Nitpick comments (2)
cmuxTests/TerminalAndGhosttyTests.swift (1)

2993-2994: Make the async settle step deterministic to reduce CI flakes.

The fixed RunLoop delay can intermittently fail under load. Prefer the existing drainMainQueue() helper so the test waits for queued main-thread work instead of wall-clock timing.

♻️ Suggested test stabilization
-        RunLoop.current.run(until: Date().addingTimeInterval(0.05))
+        drainMainQueue()
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@cmuxTests/TerminalAndGhosttyTests.swift` around lines 2993 - 2994, The test
currently uses a fixed delay via RunLoop.current.run(until:
Date().addingTimeInterval(0.05)) which is flaky; replace that line with a call
to the existing drainMainQueue() helper so the test deterministically waits for
queued main-thread work to complete (i.e., remove the RunLoop.current.run(...)
invocation and call drainMainQueue() in its place, ensuring drainMainQueue() is
in scope for the test).
.github/workflows/ci.yml (1)

510-520: Make the flake detector file-based and newline-safe.

Line 513's predicate only matches if both fragments land on the same log line. grep won't let .* span newlines, so the retry can be skipped if XCTest wraps that activation error. Grepping the file directly also removes the unnecessary OUTPUT=$(cat ...) copies.

♻️ Proposed change
-          OUTPUT=$(cat /tmp/display-resolution-ui-test-output.txt)
           set -e

-          if [ "$EXIT_CODE" -ne 0 ] && echo "$OUTPUT" | grep -q "Failed to activate application.*Running Background"; then
+          if [ "$EXIT_CODE" -ne 0 ] \
+            && grep -q "Failed to activate application" /tmp/display-resolution-ui-test-output.txt \
+            && grep -q "Running Background" /tmp/display-resolution-ui-test-output.txt; then
             echo "Display resolution UI regression hit launch activation flake, retrying once"
             pkill -x "cmux DEV" || true
             sleep 2
             set +e
             run_display_resolution_ui_test | tee /tmp/display-resolution-ui-test-output.txt
             EXIT_CODE=${PIPESTATUS[0]}
-            OUTPUT=$(cat /tmp/display-resolution-ui-test-output.txt)
             set -e
           fi
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In @.github/workflows/ci.yml around lines 510 - 520, Change the flake detector
to operate on the test output file itself and make the match newline-safe:
instead of piping the shell variable OUTPUT into grep, run a newline-aware
search like grep -qz "Failed to activate application.*Running Background"
/tmp/display-resolution-ui-test-output.txt (or use perl -0777 -ne 'print if
/Failed to activate application.*Running Background/' for portability), remove
the redundant OUTPUT=$(cat /tmp/display-resolution-ui-test-output.txt) copies
and use the file path /tmp/display-resolution-ui-test-output.txt and EXIT_CODE
for the retry predicate so the retry triggers even when the two fragments are on
different lines.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@CLI/cmux.swift`:
- Around line 10929-10969: The generateTabTitle(from:) function currently builds
the frequency map from raw messages, causing sensitive data (paths, emails,
UUIDs, tokens, ENV_VAR=..., long numbers) to appear in tab titles; before
tokenization, run the same redaction used by the Claude tag-extraction (apply
the unanchored sensitiveSpanPatterns across the combined subtitle+body string to
remove/replace sensitive spans), then proceed to tokenize the redacted combined
text, and additionally filter each resulting token against the anchored
sensitivePatterns (and skip any matches) before updating freq/casing; reference
generateTabTitle, combined, tokens, freq, and the existing
sensitiveSpanPatterns/sensitivePatterns utilities when implementing this change.
- Around line 10263-10272: The auto-titling currently sends a regular tab.action
"rename" which gets persisted as a user/custom title and can be overwritten by
later transcript state; instead send a distinct, non-persistent signal so manual
renames aren't clobbered: in the block that calls client.sendV2 (around
generateTabTitle and transcript.hasCustomTitle), replace or augment the "rename"
action with an explicit auto-title marker (e.g., use action "auto_rename" or add
a parameter like "auto_generated": true) so the app can treat it as
non-persistent; keep the transcript.hasCustomTitle guard in place and ensure the
app-side handler for "auto_rename" does not write to the same custom-title
persistence path used by manual rename actions.
- Around line 10265-10266: The variable `transcript` is undefined where it's
being checked; update the flow so the stop/idle branch has access to a
TranscriptSummary instance: either change summarizeClaudeHookStop(...) to return
the TranscriptSummary (not just (subtitle: String, body: String)?) and propagate
that value through runClaudeHook(...) so you can read transcript.hasCustomTitle
where generateTabTitle is called, or bind/obtain the TranscriptSummary directly
in the stop/idle branch before the if-check; specifically modify
summarizeClaudeHookStop, runClaudeHook, and the call site so a TranscriptSummary
(with hasCustomTitle) is available to evaluate and pass into generateTabTitle.
- Around line 10950-10952: The tokenizer's regex (pattern) only matches Han
characters and Latin words, causing Japanese and Korean input to produce no
tokens and an empty frequency dictionary; update the NSRegularExpression used in
the tokenization (the pattern variable in CLI/cmux.swift) to include Hiragana
(\u3040-\u309f), Katakana (\u30a0-\u30ff) and Hangul (\uac00-\ud7af) ranges for
multi-character runs (e.g. use ranges like
[\u4e00-\u9fff\u3040-\u309f\u30a0-\u30ff\uac00-\ud7af]{2,}) or switch to Unicode
script properties (e.g. \p{Han}, \p{Hiragana}, \p{Katakana}, \p{Hangul}) so
tokens and the frequency dictionary are correctly populated instead of returning
nil.

In `@Sources/AppDelegate.swift`:
- Around line 2425-2438: The stabilizeUITestLaunchWindowAndForeground retry loop
is repeatedly calling moveUITestWindowToTargetDisplayIfNeeded() (which itself
schedules retries), causing overlapping timers; to fix, ensure
moveUITestWindowToTargetDisplayIfNeeded() (and optionally
activateUITestAppIfNeeded()) are invoked once before starting the retry loop in
stabilizeUITestLaunchWindowAndForeground (or add a guard inside
moveUITestWindowToTargetDisplayIfNeeded to no-op if a retry is already
scheduled), and keep writeUITestDiagnosticsIfNeeded(stage:) within the loop for
per-attempt logging so you avoid nested retry storms while preserving
stabilization retries.

---

Nitpick comments:
In @.github/workflows/ci.yml:
- Around line 510-520: Change the flake detector to operate on the test output
file itself and make the match newline-safe: instead of piping the shell
variable OUTPUT into grep, run a newline-aware search like grep -qz "Failed to
activate application.*Running Background"
/tmp/display-resolution-ui-test-output.txt (or use perl -0777 -ne 'print if
/Failed to activate application.*Running Background/' for portability), remove
the redundant OUTPUT=$(cat /tmp/display-resolution-ui-test-output.txt) copies
and use the file path /tmp/display-resolution-ui-test-output.txt and EXIT_CODE
for the retry predicate so the retry triggers even when the two fragments are on
different lines.

In `@cmuxTests/TerminalAndGhosttyTests.swift`:
- Around line 2993-2994: The test currently uses a fixed delay via
RunLoop.current.run(until: Date().addingTimeInterval(0.05)) which is flaky;
replace that line with a call to the existing drainMainQueue() helper so the
test deterministically waits for queued main-thread work to complete (i.e.,
remove the RunLoop.current.run(...) invocation and call drainMainQueue() in its
place, ensuring drainMainQueue() is in scope for the test).

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: c2d8771a-0169-464b-bd7d-669efae1c0f8

📥 Commits

Reviewing files that changed from the base of the PR and between 22c50a4 and 275ee5c.

📒 Files selected for processing (5)
  • .github/workflows/ci.yml
  • CLI/cmux.swift
  • Sources/AppDelegate.swift
  • Sources/TerminalWindowPortal.swift
  • cmuxTests/TerminalAndGhosttyTests.swift

Comment thread CLI/cmux.swift
Comment on lines +10263 to +10272
// Auto-name the tab based on conversation content.
// Skip if user has set a custom title via /rename.
if let transcript, !transcript.hasCustomTitle,
let tabTitle = generateTabTitle(from: transcript) {
_ = try? client.sendV2(method: "tab.action", params: [
"workspace_id": workspaceId,
"surface_id": surfaceId,
"action": "rename",
"title": tabTitle,
])

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Don't persist auto-titles through the manual rename action.

This only honors transcript-side custom-title events. A title set from cmux itself (rename-tab, tab context menu, etc.) is not in the Claude transcript, and tab.action "rename" stores through the same custom-title path, so the next stop will overwrite that manual title. This needs a separate auto-title path or an app-side “skip if user-customized” guard.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@CLI/cmux.swift` around lines 10263 - 10272, The auto-titling currently sends
a regular tab.action "rename" which gets persisted as a user/custom title and
can be overwritten by later transcript state; instead send a distinct,
non-persistent signal so manual renames aren't clobbered: in the block that
calls client.sendV2 (around generateTabTitle and transcript.hasCustomTitle),
replace or augment the "rename" action with an explicit auto-title marker (e.g.,
use action "auto_rename" or add a parameter like "auto_generated": true) so the
app can treat it as non-persistent; keep the transcript.hasCustomTitle guard in
place and ensure the app-side handler for "auto_rename" does not write to the
same custom-title persistence path used by manual rename actions.

Comment thread CLI/cmux.swift
Comment on lines +10265 to +10266
if let transcript, !transcript.hasCustomTitle,
let tabTitle = generateTabTitle(from: transcript) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

sed -n '10239,10273p' CLI/cmux.swift
printf '\nBindings/usages named transcript:\n'
rg -n '\b(let|var)\s+transcript\b|\bif let transcript\b' CLI/cmux.swift -C2

Repository: manaflow-ai/cmux

Length of output: 2448


🏁 Script executed:

#!/bin/bash
set -euo pipefail

# Get the function definition boundaries for runClaudeHook
rg -n 'func runClaudeHook\(' CLI/cmux.swift | head -1

# Extract a wider range around line 10265 to see the full context
echo "=== Full context from 10200 to 10290 ==="
sed -n '10200,10290p' CLI/cmux.swift

# Search for where transcript might be bound in this function
echo -e "\n=== All transcript bindings in broader context ==="
sed -n '10200,10290p' CLI/cmux.swift | rg -n 'transcript'

Repository: manaflow-ai/cmux

Length of output: 4522


🏁 Script executed:

#!/bin/bash
set -euo pipefail

# Check the type of 'completion' return from summarizeClaudeHookStop
echo "=== summarizeClaudeHookStop return type and implementation ==="
rg -n 'func summarizeClaudeHookStop' CLI/cmux.swift -A 50 | head -60

Repository: manaflow-ai/cmux

Length of output: 2485


transcript is undefined in this scope and will cause a compile error.

runClaudeHook(...) receives only a (subtitle: String, body: String)? return from summarizeClaudeHookStop(...), not the TranscriptSummary with hasCustomTitle property. The transcript variable exists only within summarizeClaudeHookStop at line 10831 and is never returned. Either return TranscriptSummary from summarizeClaudeHookStop, or bind transcript directly in the stop/idle branch before using it at lines 10265–10266.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@CLI/cmux.swift` around lines 10265 - 10266, The variable `transcript` is
undefined where it's being checked; update the flow so the stop/idle branch has
access to a TranscriptSummary instance: either change
summarizeClaudeHookStop(...) to return the TranscriptSummary (not just
(subtitle: String, body: String)?) and propagate that value through
runClaudeHook(...) so you can read transcript.hasCustomTitle where
generateTabTitle is called, or bind/obtain the TranscriptSummary directly in the
stop/idle branch before the if-check; specifically modify
summarizeClaudeHookStop, runClaudeHook, and the call site so a TranscriptSummary
(with hasCustomTitle) is available to evaluate and pass into generateTabTitle.

Comment thread CLI/cmux.swift
Comment on lines +10929 to +10969
private func generateTabTitle(from transcript: TranscriptSummary) -> String? {
let allMessages = transcript.firstUserMessages + transcript.lastUserMessages
guard !allMessages.isEmpty else { return nil }

let combined = allMessages.joined(separator: " ")

// Stop words (English + Chinese) to filter out
let stopWords: Set<String> = [
"please", "can", "you", "the", "a", "an", "i", "want", "to", "me",
"need", "help", "with", "this", "that", "for", "in", "on", "it",
"is", "my", "do", "let", "make", "also", "and", "or", "but",
"just", "some", "all", "any", "be", "have", "will", "would",
"should", "could", "from", "how", "what", "when", "where", "which",
"not", "no", "yes", "use", "code", "file", "run", "test",
"请", "帮我", "帮", "我", "你", "能", "可以", "一下", "看看",
"这个", "那个", "的", "了", "吗", "呢", "是", "在", "有", "和",
"把", "给", "让", "对", "到", "从", "用", "也", "都", "就",
"还", "又", "很", "不", "没", "怎么", "什么", "如何", "已经",
"然后", "现在", "应该", "可能", "需要", "想", "要", "会",
]

// Tokenize: CJK runs (2+ chars) or Latin word tokens
var tokens: [String] = []
let pattern = try? NSRegularExpression(pattern: "[\\u4e00-\\u9fff]{2,}|[a-zA-Z][a-zA-Z0-9]*", options: [])
let nsRange = NSRange(combined.startIndex..., in: combined)
pattern?.enumerateMatches(in: combined, range: nsRange) { match, _, _ in
guard let match, let range = Range(match.range, in: combined) else { return }
tokens.append(String(combined[range]))
}

// Count frequencies, excluding stop words
var freq: [String: Int] = [:]
var casing: [String: String] = [:]
for token in tokens {
let low = token.lowercased()
guard !stopWords.contains(low), low.count >= 2 else { continue }
freq[low, default: 0] += 1
if casing[low] == nil {
casing[low] = token
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Redact sensitive spans before counting keywords.

This frequency map is built from raw user prompts, so paths, emails, UUIDs, access tokens, ENV_VAR=..., and long numerics can surface directly as visible tab titles. Reuse the existing Claude tag-extraction redaction/token filters before tokenization.

Based on learnings, pre-redact sensitive spans (UUIDs, emails, access tokens, filesystem paths, ENV_VAR=..., long numerics) across the combined body+subtitle using unanchored sensitiveSpanPatterns before tokenization. Then tokenize and still filter each token with anchored sensitivePatterns.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@CLI/cmux.swift` around lines 10929 - 10969, The generateTabTitle(from:)
function currently builds the frequency map from raw messages, causing sensitive
data (paths, emails, UUIDs, tokens, ENV_VAR=..., long numbers) to appear in tab
titles; before tokenization, run the same redaction used by the Claude
tag-extraction (apply the unanchored sensitiveSpanPatterns across the combined
subtitle+body string to remove/replace sensitive spans), then proceed to
tokenize the redacted combined text, and additionally filter each resulting
token against the anchored sensitivePatterns (and skip any matches) before
updating freq/casing; reference generateTabTitle, combined, tokens, freq, and
the existing sensitiveSpanPatterns/sensitivePatterns utilities when implementing
this change.

Comment thread CLI/cmux.swift
Comment on lines +10950 to +10952
// Tokenize: CJK runs (2+ chars) or Latin word tokens
var tokens: [String] = []
let pattern = try? NSRegularExpression(pattern: "[\\u4e00-\\u9fff]{2,}|[a-zA-Z][a-zA-Z0-9]*", options: [])

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
python - <<'PY'
import re

pat = re.compile(r'[\u4e00-\u9fff]{2,}|[a-zA-Z][a-zA-Z0-9]*')
samples = [
    "请修复登录问题",
    "このテストを直して",
    "버그 수정해줘",
]

for s in samples:
    print(f"{s!r} -> {pat.findall(s)}")
PY

Repository: manaflow-ai/cmux

Length of output: 118


🏁 Script executed:

cd /tmp && git clone --depth 1 https://github.com/manaflow-ai/cmux.git 2>&1 | head -20

Repository: manaflow-ai/cmux

Length of output: 910


🏁 Script executed:

wc -l CLI/cmux.swift

Repository: manaflow-ai/cmux

Length of output: 80


🏁 Script executed:

sed -n '10940,10970p' CLI/cmux.swift

Repository: manaflow-ai/cmux

Length of output: 1526


🏁 Script executed:

sed -n '10900,10950p' CLI/cmux.swift

Repository: manaflow-ai/cmux

Length of output: 2593


🏁 Script executed:

sed -n '10960,10990p' CLI/cmux.swift

Repository: manaflow-ai/cmux

Length of output: 1088


The tokenizer misses Hiragana, Katakana, and Hangul, causing silent failure for Japanese and Korean prompts.

The regex [\u4e00-\u9fff]{2,}|[a-zA-Z][a-zA-Z0-9]* only matches Han ideographs and Latin words. Japanese prompts like このテストを直して and Korean prompts like 버그 수정해줘 produce no tokens, resulting in an empty frequency dictionary that triggers a silent nil return at line 10968. Widen the regex to include Hiragana (\u3040-\u309f), Katakana (\u30a0-\u30ff), and Hangul (\uac00-\ud7af) ranges, or switch to Unicode script properties.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@CLI/cmux.swift` around lines 10950 - 10952, The tokenizer's regex (pattern)
only matches Han characters and Latin words, causing Japanese and Korean input
to produce no tokens and an empty frequency dictionary; update the
NSRegularExpression used in the tokenization (the pattern variable in
CLI/cmux.swift) to include Hiragana (\u3040-\u309f), Katakana (\u30a0-\u30ff)
and Hangul (\uac00-\ud7af) ranges for multi-character runs (e.g. use ranges like
[\u4e00-\u9fff\u3040-\u309f\u30a0-\u30ff\uac00-\ud7af]{2,}) or switch to Unicode
script properties (e.g. \p{Han}, \p{Hiragana}, \p{Katakana}, \p{Hangul}) so
tokens and the frequency dictionary are correctly populated instead of returning
nil.

Comment thread Sources/AppDelegate.swift
Comment on lines +2425 to +2438
moveUITestWindowToTargetDisplayIfNeeded()
activateUITestAppIfNeeded()

let hasWindow = !NSApp.windows.isEmpty
let hasVisibleWindow = NSApp.windows.contains { $0.isVisible }
let hasKeyWindow = NSApp.keyWindow != nil
let stage = attempt == 0 ? "afterForceWindow" : "afterForceWindow.retry\(attempt)"
writeUITestDiagnosticsIfNeeded(stage: stage)

guard attempt < 20 else { return }
if !hasWindow || !hasVisibleWindow || !hasKeyWindow || !NSRunningApplication.current.isActive {
DispatchQueue.main.asyncAfter(deadline: .now() + 0.25) { [weak self] in
self?.stabilizeUITestLaunchWindowAndForeground(attempt: attempt + 1)
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Avoid overlapping retry loops during UI-test stabilization

This retry loop calls moveUITestWindowToTargetDisplayIfNeeded() on every attempt, but that helper also schedules its own retries. The nested retries can create timer storms and flaky timing/diagnostics under CI load.

💡 Proposed fix
-        moveUITestWindowToTargetDisplayIfNeeded()
+        moveUITestWindowToTargetDisplayIfNeeded(attempt: attempt, allowRetry: false)
         activateUITestAppIfNeeded()
-    private func moveUITestWindowToTargetDisplayIfNeeded(attempt: Int = 0) {
+    private func moveUITestWindowToTargetDisplayIfNeeded(
+        attempt: Int = 0,
+        allowRetry: Bool = true
+    ) {
         let env = ProcessInfo.processInfo.environment
         guard let rawDisplayID = env["CMUX_UI_TEST_TARGET_DISPLAY_ID"],
               let targetDisplayID = UInt32(rawDisplayID) else {
             return
         }

         guard let screen = NSScreen.screens.first(where: { $0.cmuxDisplayID == targetDisplayID }) else {
-            if attempt < 20 {
+            if allowRetry, attempt < 20 {
                 DispatchQueue.main.asyncAfter(deadline: .now() + 0.25) { [weak self] in
-                    self?.moveUITestWindowToTargetDisplayIfNeeded(attempt: attempt + 1)
+                    self?.moveUITestWindowToTargetDisplayIfNeeded(
+                        attempt: attempt + 1,
+                        allowRetry: allowRetry
+                    )
                 }
             }
             self.writeUITestDiagnosticsIfNeeded(stage: "targetDisplayMissing")
             return
         }

         guard let window = NSApp.windows.first else {
-            if attempt < 20 {
+            if allowRetry, attempt < 20 {
                 DispatchQueue.main.asyncAfter(deadline: .now() + 0.25) { [weak self] in
-                    self?.moveUITestWindowToTargetDisplayIfNeeded(attempt: attempt + 1)
+                    self?.moveUITestWindowToTargetDisplayIfNeeded(
+                        attempt: attempt + 1,
+                        allowRetry: allowRetry
+                    )
                 }
             }
             self.writeUITestDiagnosticsIfNeeded(stage: "targetDisplayNoWindow")
             return
         }

         // ...
-        if window.screen?.cmuxDisplayID != targetDisplayID, attempt < 20 {
+        if window.screen?.cmuxDisplayID != targetDisplayID, allowRetry, attempt < 20 {
             DispatchQueue.main.asyncAfter(deadline: .now() + 0.25) { [weak self] in
-                self?.moveUITestWindowToTargetDisplayIfNeeded(attempt: attempt + 1)
+                self?.moveUITestWindowToTargetDisplayIfNeeded(
+                    attempt: attempt + 1,
+                    allowRetry: allowRetry
+                )
             }
             return
         }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@Sources/AppDelegate.swift` around lines 2425 - 2438, The
stabilizeUITestLaunchWindowAndForeground retry loop is repeatedly calling
moveUITestWindowToTargetDisplayIfNeeded() (which itself schedules retries),
causing overlapping timers; to fix, ensure
moveUITestWindowToTargetDisplayIfNeeded() (and optionally
activateUITestAppIfNeeded()) are invoked once before starting the retry loop in
stabilizeUITestLaunchWindowAndForeground (or add a guard inside
moveUITestWindowToTargetDisplayIfNeeded to no-op if a retry is already
scheduled), and keep writeUITestDiagnosticsIfNeeded(stage:) within the loop for
per-attempt logging so you avoid nested retry storms while preserving
stabilization retries.

@Horacehxw

Copy link
Copy Markdown
Author

Superseded by #2043 which uses AI-powered summaries instead of keyword extraction.

@Horacehxw Horacehxw closed this Mar 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants