Skip to content

iOS: run voice-dictation audio activation off the main thread (fix mic-button animation lag) - #6868

Merged
austinywang merged 6 commits into
mainfrom
issue-6284-ios-voice-dictation-mic-button-has
Jun 29, 2026
Merged

austinywang merged 6 commits into
mainfrom
issue-6284-ios-voice-dictation-mic-button-has

Conversation

@austinywang

@austinywang austinywang commented Jun 26, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #6284

Problem

On iOS, tapping the composer mic button to start/stop voice dictation produced a visible animation hitch. The whole dictation start/stop path ran synchronously on the @MainActor ComposerDictationController:

DICT M begin.session.setActive
DICT M begin.engine.start

AVAudioSession.setActive(true) and AVAudioEngine.start() (and the symmetric engine.stop() + setActive(false) on stop) are synchronous audio-hardware calls that block the caller ~100-300ms each, freezing the mic button / field animation on every press.

The crash and text-loss bugs were already fixed separately; this PR tackles the remaining animation lag, via the architectural change the issue proposed.

Fix

Extract a thread-safe ComposerDictationAudioEngine (@unchecked Sendable) that owns the AVAudioEngine + shared AVAudioSession lifecycle on its own serial queue, exposing start(tapBlock:onReady:) / stop() with @Sendable callbacks. The main actor now only enqueues the activation/teardown — it never blocks on the audio hardware.

  • The tap block captures the non-Sendable SFSpeechAudioBufferRecognitionRequest via nonisolated(unsafe) (its append is thread-safe), so it can cross into the owner's queue.
  • The non-Sendable request / recognizer never leave the main actor — the recognition task is created in the main-actor engine-ready callback (handleEngineReady), sidestepping the Swift 6 region-isolation errors the issue notes a naive Task.detached hits.

Supersession (new async window)

Because activation is now asynchronous, a second mic tap / send / navigation during the ~100-300ms spin-up can abandon the start. A monotonic startToken (bumped on every start and every teardown) plus a pure, host-testable helper composerDictationStartDisposition(callbackToken:currentToken:state:) let a late engine-ready callback detect it was superseded and discard its result — preventing a double-started engine or a leaked input tap. teardown()'s off-main stop() is serialized on the owner's queue after the in-flight start's activation and before any later start, so the engine is reliably torn down and never double-started.

Files

  • ComposerDictationAudioEngine.swift (new) — off-main @unchecked Sendable audio engine owner.
  • ComposerDictationController.swift — delegate activation/teardown to the owner; add startToken + handleEngineReady; drop didActivateSession/stopEngineAndSession (now internal to the owner).
  • ComposerDictationTextMerger.swift — add the pure composerDictationStartDisposition supersession helper (host-compilable, no Speech/AVFoundation).
  • ComposerDictationTests.swift — unit tests for the supersession rule.

Testing

  • Host unit tests: 68 pass, including 5 new startDisposition* tests covering the supersession partition (token match × state).
  • iOS type-check: the whole CmuxMobileSupport module type-checks against the iOS 26.2 SDK in Swift 6 language mode with no new errors/warnings (the @unchecked Sendable owner, the @Sendable closures crossing the queue, and the nonisolated(unsafe) capture all compile).
  • The core threading fix (off-main audio) is inherently a timing/animation change and is not cleanly host-testable (iOS-only AVAudioEngine, requires a device + mic permission, timing-flaky), so it's verified by behavior on-device rather than an automated regression test; the new supersession logic that the async change introduces is unit-tested.
  • Localization: no user-facing strings added or changed (pure internal threading/architecture change).
  • No public API change: toggle/stop/cancel/state/isAvailable/locksComposerField are unchanged, so the other consumer (ChatComposerView) is unaffected.

🤖 Generated with Claude Code


Summary by cubic

Runs voice dictation audio activation off the main thread to eliminate mic-button animation hitches, and locks the composer field during engine spin-up to prevent edit loss. Fixes #6284.

  • Bug Fixes

    • Moved AVAudioSession.setActive and AVAudioEngine.start/stop to a background serial queue to remove 100–300ms UI freezes when tapping the mic.
    • Locked the composer field from .requestingPermission through .listening/.stopping to avoid overwriting text typed during async engine spin-up.
    • Added a monotonic start token and logic to discard stale engine-ready callbacks, preventing double-starts and leaked input taps.
  • Refactors

    • Introduced ComposerDictationAudioEngine with start(tapBlock:onReady:) and stop(), and updated the controller to create the recognition task on engine-ready (main actor).
    • Scoped the supersession helper to ComposerDictationState.startDisposition(callbackToken:currentToken:); added unit tests for this rule and the field-lock behavior.
    • Documented and enforced the audio-engine serial-queue carve-out; added dispatchPrecondition(.onQueue(...)) in teardown to guard isolation.
    • Satisfied lint/policy checks: moved start-disposition types to ComposerDictationStartDisposition.swift, made makeTapBlock a nonisolated instance method, and added a nearby safety note for the nonisolated(unsafe) capture (no behavior change).

Written for commit c3e95f5. Summary will update on new commits.

Review in cubic

Summary by CodeRabbit

  • Bug Fixes
    • Improved dictation start/stop reliability, especially when starting and stopping quickly.
    • Reduced missed or delayed dictation responses by handling audio setup off the main thread.
    • Prevented stale dictation callbacks from affecting a newer recording attempt.
    • Made audio cleanup more dependable so microphone access and session state are released correctly after errors or cancellations.

The composer mic button hitched on every press because the whole
dictation start/stop path ran synchronously on the @mainactor
ComposerDictationController: AVAudioSession.setActive(true) and
AVAudioEngine.start() (and their stop counterparts) are blocking
audio-hardware calls (~100-300ms each) that froze the button/field
animation (issue #6284).

Extract a thread-safe ComposerDictationAudioEngine (@unchecked Sendable)
that owns the AVAudioEngine + shared AVAudioSession lifecycle on its own
serial queue, exposing start(tapBlock:onReady:)/stop() with @sendable
callbacks, so the main actor only ever enqueues the work and never
blocks on the hardware. The tap block captures the non-Sendable
SFSpeechAudioBufferRecognitionRequest via nonisolated(unsafe) (append is
thread-safe); the non-Sendable request/recognizer never leave the main
actor — the recognition task is created in the main-actor engine-ready
callback.

Because activation is now asynchronous, a second mic tap / send /
navigation during the ~100-300ms spin-up can abandon the start. A
monotonic startToken (bumped on every start and teardown) plus the pure,
host-testable composerDictationStartDisposition(...) helper let a late
engine-ready callback detect it was superseded and discard its result,
preventing a double-started engine or leaked tap. teardown()'s off-main
stop is serialized after the in-flight start and before any later start
on the owner's queue.

No user-facing strings changed. The threading fix is iOS-only and not
host-testable; the new supersession logic is covered by host unit tests,
and the module type-checks against the iOS 26.2 SDK in Swift 6 mode.

Fixes #6284

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@vercel

vercel Bot commented Jun 26, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
cmux Canceled Canceled Jun 26, 2026 3:19pm
cmux-staging Building Building Preview, Comment Jun 26, 2026 3:19pm

@coderabbitai

coderabbitai Bot commented Jun 26, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Dictation now starts audio capture through a dedicated iOS audio engine helper, routes engine-ready callbacks through a token/state disposition helper, and updates controller start, stop, and teardown paths to use the new off-main lifecycle.

Changes

Dictation audio lifecycle

Layer / File(s) Summary
Audio engine owner
Packages/iOS/CmuxMobileSupport/Sources/CmuxMobileSupport/ComposerDictationAudioEngine.swift
Defines a serial-queue ComposerDictationAudioEngine that configures AVAudioSession, installs the input tap, starts AVAudioEngine, and tears down session and tap state.
Callback disposition contract and tests
Packages/iOS/CmuxMobileSupport/Sources/CmuxMobileSupport/ComposerDictationTextMerger.swift, Packages/iOS/CmuxMobileSupport/Tests/CmuxMobileSupportTests/ComposerDictationTests.swift
Adds composerDictationStartDisposition(...) and tests that engine-ready callbacks are applied only for a matching token while dictation is still requesting permission.
Start handoff and supersession
Packages/iOS/CmuxMobileSupport/Sources/CmuxMobileSupport/ComposerDictationController.swift
Replaces controller-owned audio setup with ComposerDictationAudioEngine, tracks each start with startToken, routes pending-start cancellation through teardown(), and builds the tap block for off-main execution.
Stop and teardown
Packages/iOS/CmuxMobileSupport/Sources/CmuxMobileSupport/ComposerDictationController.swift
Updates stop() and teardown() to end the request, invalidate in-flight starts, clear recognition state, and stop the audio engine owner instead of directly managing session shutdown.

Sequence Diagram(s)

sequenceDiagram
  participant ComposerDictationController
  participant ComposerDictationAudioEngine
  participant AVAudioSession
  participant AVAudioEngine
  participant composerDictationStartDisposition

  ComposerDictationController->>ComposerDictationAudioEngine: start(tapBlock:onReady:)
  ComposerDictationAudioEngine->>AVAudioSession: setCategory(.record, .measurement)
  ComposerDictationAudioEngine->>AVAudioSession: setActive(true)
  ComposerDictationAudioEngine->>AVAudioEngine: installTap() / prepare() / start()
  ComposerDictationAudioEngine-->>ComposerDictationController: onReady(true)
  ComposerDictationController->>composerDictationStartDisposition: evaluate callbackToken/currentToken/state
  composerDictationStartDisposition-->>ComposerDictationController: .apply or .discardStale
  ComposerDictationController->>ComposerDictationAudioEngine: stop()
  ComposerDictationAudioEngine->>AVAudioEngine: stop() / removeTap()
  ComposerDictationAudioEngine->>AVAudioSession: setActive(false, .notifyOthersOnDeactivation)
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related issues

Possibly related PRs

  • manaflow-ai/cmux#6290: Both changes rework dictation teardown and session ownership around whether dictation actually activated the session.

Poem

I hopped by the mic with a twitchy nose,
Off-main the audio engine rose.
Tokens danced and stale ones fled,
Permission patted, then words were fed.
🐇🎤 Carrots for calm; dictation goes!


Important

Pre-merge checks failed

Please resolve all errors before merging. Addressing warnings is optional.

❌ Failed checks (2 errors, 2 warnings)

Check name Status Explanation Resolution
Cmux Swift Blocking Runtime ❌ Error ComposerDictationController.stop() adds a production Task.sleep watchdog before finishGraceful(), which is forbidden timing-based synchronization. Replace the sleep-based watchdog with a real completion/signal from recognition, or a cancellation-aware timer abstraction; keep any sleeps test-only.
Cmux No Ambient Global State ❌ Error ComposerDictationTextMerger.swift:102 adds an internal file-scope free function API; the rule forbids new ambient global helpers instead of owned methods. Move the disposition check onto ComposerDictationController as a private helper, or make it private/fileprivate if it’s only a local utility.
Docstring Coverage ⚠️ Warning Docstring coverage is 72.22% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
Description check ⚠️ Warning The description is detailed, but it does not follow the required template: it lacks the Summary heading, Demo Video, Review Trigger block, and Checklist. Add the missing template sections and include the requested review-trigger comment, demo link if applicable, and checklist items.
✅ Passed checks (21 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the main change: moving iOS voice-dictation audio activation off the main thread to fix UI lag.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Cmux Swift Actor Isolation ✅ Passed PASS: The new @MainActor controller stays UI-bound, and the new @unchecked Sendable audio-engine owner is explicitly queue-confined with documented safety.
Cmux Browser Automation Off-Main ✅ Passed The PR only changes iOS dictation engine/controller/test files; no browser.* commands, router policy, or WebKit wait paths are touched.
Cmux Expensive Synchronous Load ✅ Passed PASS: The PR only changes iOS dictation audio/recognition code; no agent-history/JSON/transcript loaders or main-actor sync scans were added or moved.
Cmux Cache Substitution Correctness ✅ Passed No fresh authoritative read was replaced by a cached value; this is an off-main audio refactor with a token-based freshness check, not a persistence/history/snapshot path.
Cmux No Hacky Sleeps ✅ Passed PR only changes Swift sources; the lone Task.sleep is a Swift watchdog, and no non-Swift runtime sleep hacks were introduced.
Cmux Algorithmic Complexity ✅ Passed Production code only adds O(1) state/engine logic; loops are confined to small tests and no scalable collections are rescanned.
Cmux Swift Concurrency ✅ Passed The new DispatchQueue and callbacks are confined to AVFoundation/Speech OS boundaries; no new Combine state or lifecycle-bearing fire-and-forget Tasks were added.
Cmux Swift @Concurrent ✅ Passed No changed Swift code adds misused @concurrent; async callbacks explicitly hop to @MainActor, and blocking audio work is enqueued on a serial queue.
Cmux Swift File And Package Boundaries ✅ Passed PASS: The change stays in CmuxMobileSupport, adds a 120-line focused iOS audio owner and a 153-line pure helper, and doesn’t create oversized or mixed-responsibility files.
Cmux Swiftpm Lockfiles ✅ Passed Diff only updates vendored third-party git submodules (ghostty, vendor/bonsplit); no cmux-owned Package.swift/.gitignore/Xcode Package.resolved changes.
Cmux Swift Logging ✅ Passed No added print/debugPrint/dump/NSLog/Logger in the diff; only comments/docs changed, and tests/debug output are exempt.
Cmux User-Facing Error Privacy ✅ Passed PASS: The diff only adds internal dictation engine/state logic and tests; no new user-facing error, alert, or recovery copy was introduced.
Cmux Full Internationalization ✅ Passed Only internal code and tests changed; no user-facing text, string catalogs, Info.plist, or web locale files were added or edited.
Cmux Swiftui State Layout ✅ Passed No SwiftUI layout/state anti-patterns appear in the diff: only controller/audio-engine/test files changed, and the controller uses @Observable, not legacy store patterns.
Cmux Architecture Rethink ✅ Passed Uses an explicit serial-queue audio owner plus token/state supersession; no banned sleeps, locks, duplicate wiring, or unclear lifecycle ownership introduced.
Cmux Swift Auxiliary Window Close Shortcuts ✅ Passed Touched files are dictation audio/controller/helper/tests only; no NSWindow/NSPanel/NSWindowController/Window/WindowGroup or cmuxAuxiliaryWindowIdentifiers changes.
Cmux Source Artifacts ✅ Passed Only changed paths are submodule pointer bumps (ghostty, vendor/bonsplit); no logs, temp dirs, build output, or artifact paths appear in the diff.
Cmux No Test Or Debug Seam In Production Source ✅ Passed No new test/debug seam was added in production Sources; the new helper has a real controller caller and test observation stays in Tests/ via @testable import.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch issue-6284-ios-voice-dictation-mic-button-has

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@greptile-apps

greptile-apps Bot commented Jun 26, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR extracts ComposerDictationAudioEngine to move AVAudioSession.setActive and AVAudioEngine.start/stop (each ~100–300ms synchronous hardware calls) off the @MainActor ComposerDictationController onto a dedicated private serial DispatchQueue, eliminating the mic-button animation hitch on tap. A monotonic startToken plus ComposerDictationState.startDisposition(callbackToken:currentToken:) handle the new async window between enqueuing activation and the engine-ready callback arriving.

  • ComposerDictationAudioEngine (new): owns the AVAudioEngine + AVAudioSession lifecycle on a private serial queue, with start(tapBlock:onReady:) / stop() serialized in FIFO order so a rapid stop-during-spin-up always tears down before any later start.
  • ComposerDictationController: delegates audio work to the new owner, replaces didActivateSession with a startToken counter, adds handleEngineReady to create the recognition task on the main actor, and extends the field-lock to cover .requestingPermission to close the async edit-loss window.
  • ComposerDictationStartDisposition + tests: pure supersession helper scoped as a method on ComposerDictationState, exhaustively unit-tested across the full token × state partition (80 combinations).

Confidence Score: 5/5

Safe to merge. The threading invariants are sound, the supersession logic is exhaustively unit-tested, and the change is strictly additive to the existing public API surface.

The audio engine carve-out correctly confines all blocking hardware calls to a private serial queue; dispatchPrecondition enforces the isolation contract at runtime. The monotonic startToken plus state-gated startDisposition form a correct supersession check: both the token comparison and the requestingPermission state guard must hold before a callback can transition to .listening. FIFO ordering of stop() after start() on the same serial queue eliminates double-start and tap-leak scenarios. The nonisolated(unsafe) weak capture of the non-Sendable request is safe because SFSpeechAudioBufferRecognitionRequest.append(_:) is documented thread-safe. No new user-facing strings, no test seams in production source, no ambient global state.

No files require special attention.

Important Files Changed

Filename Overview
Packages/iOS/CmuxMobileSupport/Sources/CmuxMobileSupport/ComposerDictationAudioEngine.swift New file. Well-isolated serial-queue audio owner with clear invariant documentation, dispatchPrecondition guard, idempotent teardown, and correct isActive gating to avoid spurious session touches.
Packages/iOS/CmuxMobileSupport/Sources/CmuxMobileSupport/ComposerDictationController.swift Delegates audio lifecycle to the new engine owner; adds startToken supersession and handleEngineReady; correct weak-capture and main-actor hop; all state transitions remain consistent.
Packages/iOS/CmuxMobileSupport/Sources/CmuxMobileSupport/ComposerDictationStartDisposition.swift New file. Clean return-type enum with cases (not a namespace), and startDisposition correctly scoped as an instance method on ComposerDictationState — addresses the pattern flagged in the previous review thread.
Packages/iOS/CmuxMobileSupport/Sources/CmuxMobileSupport/ComposerDictationTextMerger.swift Extends locksComposerField to include .requestingPermission, closing the async edit-loss window introduced by the off-main engine spin-up. No logic change to the text merger itself.
Packages/iOS/CmuxMobileSupport/Tests/CmuxMobileSupportTests/ComposerDictationTests.swift Adds 5 new supersession tests covering the full token x state partition; updates field-lock test to match the extended locksComposerField (now includes .requestingPermission).

Sequence Diagram

%%{init: {'theme': 'neutral'}}%%
sequenceDiagram
    participant U as User (mic tap)
    participant C as ComposerDictationController (@MainActor)
    participant AE as ComposerDictationAudioEngine (serial queue)
    participant SR as SFSpeechRecognizer (arbitrary queue)

    U->>C: toggle()
    C->>C: "state = .requestingPermission, startToken += 1"
    C->>AE: start(tapBlock:onReady:) [async enqueue]
    Note over C: returns immediately, no block

    AE->>AE: setActive(true), installTap(), engine.start()
    AE-->>C: onReady(true) [on audio queue]
    C->>C: "Task @MainActor enqueued"

    alt "token matches and state == .requestingPermission"
        C->>SR: recognitionTask(with: request)
        C->>C: "state = .listening"
        SR-->>C: "partials -> onText(merged text)"
    else superseded by 2nd tap, send, or nav
        C->>C: "teardown(): startToken += 1, audioEngine.stop() enqueued"
        C->>C: handleEngineReady: discardStale, return
        Note over AE: stop() serialized after start() on FIFO queue
    end

    U->>C: toggle() stop
    C->>C: "state = .stopping"
    C->>AE: stop() [async enqueue]
    AE->>AE: engine.stop(), removeTap(), setActive(false)
    SR-->>C: "final result -> finishGraceful(), state = .idle"
Loading
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
sequenceDiagram
    participant U as User (mic tap)
    participant C as ComposerDictationController (@MainActor)
    participant AE as ComposerDictationAudioEngine (serial queue)
    participant SR as SFSpeechRecognizer (arbitrary queue)

    U->>C: toggle()
    C->>C: "state = .requestingPermission, startToken += 1"
    C->>AE: start(tapBlock:onReady:) [async enqueue]
    Note over C: returns immediately, no block

    AE->>AE: setActive(true), installTap(), engine.start()
    AE-->>C: onReady(true) [on audio queue]
    C->>C: "Task @MainActor enqueued"

    alt "token matches and state == .requestingPermission"
        C->>SR: recognitionTask(with: request)
        C->>C: "state = .listening"
        SR-->>C: "partials -> onText(merged text)"
    else superseded by 2nd tap, send, or nav
        C->>C: "teardown(): startToken += 1, audioEngine.stop() enqueued"
        C->>C: handleEngineReady: discardStale, return
        Note over AE: stop() serialized after start() on FIFO queue
    end

    U->>C: toggle() stop
    C->>C: "state = .stopping"
    C->>AE: stop() [async enqueue]
    AE->>AE: engine.stop(), removeTap(), setActive(false)
    SR-->>C: "final result -> finishGraceful(), state = .idle"
Loading

Reviews (5): Last reviewed commit: "iOS: satisfy Aziz policy checks for dict..." | Re-trigger Greptile

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In
`@Packages/iOS/CmuxMobileSupport/Sources/CmuxMobileSupport/ComposerDictationController.swift`:
- Around line 398-400: The closure in ComposerDictationController.swift uses an
invalid weak reference declaration (`weak let weakRequest`), which prevents
compilation. Update the `weakRequest` binding to use a weak variable form
compatible with ARC, and keep the `nonisolated(unsafe)` capture around the
`SFSpeechAudioBufferRecognitionRequest` reference. Verify the `return { buffer,
_ in ... }` closure still appends only when the request is alive and that the
symbol `weakRequest` remains the single capture point.
- Around line 235-236: The request is being ended before the audio input tap is
fully removed, so the tap can still append audio after shutdown starts. Update
the stop/teardown flow in ComposerDictationController (the code around
request?.endAudio() and audioEngine.stop()) so endAudio() happens only after
removeTap has completed, or guard the tap callback before ending the request.
Use the existing request, audioEngine, and tap-removal sequence in
ComposerDictationController to keep the shutdown ordering correct.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 9e22a469-6e88-4e73-8118-8ab6e4d10c54

📥 Commits

Reviewing files that changed from the base of the PR and between 6d6c701 and b7d6550.

📒 Files selected for processing (4)
  • Packages/iOS/CmuxMobileSupport/Sources/CmuxMobileSupport/ComposerDictationAudioEngine.swift
  • Packages/iOS/CmuxMobileSupport/Sources/CmuxMobileSupport/ComposerDictationController.swift
  • Packages/iOS/CmuxMobileSupport/Sources/CmuxMobileSupport/ComposerDictationTextMerger.swift
  • Packages/iOS/CmuxMobileSupport/Tests/CmuxMobileSupportTests/ComposerDictationTests.swift

@blacksmith-sh

blacksmith-sh Bot commented Jun 26, 2026 •

Copy link
Copy Markdown

Found 1 test failure on Blacksmith runners:

Failure

Test View Logs
claude wrapper binary resolution checks failed/
claude wrapper binary resolution checks failed
View Logs

Fix with Codesmith
Need help on this PR? Tag /codesmith with what you need.

cmux and others added 4 commits June 26, 2026 03:13
Making engine activation asynchronous (issue #6284) opened a 100-300ms
editable window: in the already-authorized path the controller now stays
in `.requestingPermission` while the engine spins up off-main, and that
state did not set `locksComposerField`. Text typed in that window is not
in the captured `baseText`, so the first speech partial (base +
transcript) overwrote it. The previous synchronous start reached
`.listening` before returning, so the field locked immediately.

Lock the field from `.requestingPermission` through `.listening` and
`.stopping`, restoring the original instant lock. During the genuine
first-ever auth-pending flavor the system permission alert is modal, so
the field is not interactable anyway and the lock is harmless.

Found by structured review on PR #6868.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The iOS package-conventions lint (`free-function` rule) requires
functionality to be scoped to a type, not a top-level free function.
Move `composerDictationStartDisposition(callbackToken:currentToken:state:)`
to an instance method `ComposerDictationState.startDisposition(callbackToken:currentToken:)`,
alongside the enum's existing `isListening`/`locksComposerField`
accessors. Behavior is unchanged; the result enum keeps its cases (not a
namespace type), and tests/call site use the method form.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Structured review flagged the new owner's serial DispatchQueue +
@unchecked Sendable as a manual-synchronization island. Keep the queue
(it is the right tool, not an actor) but make the rationale and the
invariant explicit:

- Document why an actor is wrong here: setActive/engine.start/engine.stop
  are synchronous ~100-300ms blocking hardware calls that would block a
  cooperative-pool thread on an actor; and the supersession invariant
  needs stop() enqueued synchronously, in deterministic FIFO order, from
  the @mainactor controller's sync path — which a serial DispatchQueue
  gives and a cross-actor `await` (Task { await … }) does not. This
  mirrors the established AVFoundation-session-on-a-serial-queue pattern
  already used for capture in QRCodeCaptureController.
- Self-enforce the isolation contract with
  dispatchPrecondition(.onQueue(queue)) in teardownLocked(), so a future
  off-queue caller traps loudly instead of silently racing — directly
  addressing the "safety depends on remembering to hop through the queue"
  concern.

Carve-out marked lint:allow serial-audio-queue.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Three cmux-policy-check findings on the dictation diff:

- Add a nearby safety-argument comment at the `nonisolated(unsafe)` tap
  capture in makeTapBlock (the doc comment was >3 lines away).
- Move the added `ComposerDictationStartDisposition` enum + its
  `ComposerDictationState.startDisposition` extension into their own file
  (one major type per file), restoring ComposerDictationTextMerger.swift
  to its two pre-existing types.
- Make `makeTapBlock` a `nonisolated` instance method instead of a
  `static` one (it uses no `self`), mirroring `makeRecognitionResultHandler`
  and clearing the static-as-namespace heuristic — the enclosing
  controller is heavily stateful, so the static form was a false positive.

Behavior unchanged; host tests, iOS type-check, lint, and budget all pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@austinywang
austinywang merged commit 7ef88a6 into main Jun 29, 2026
46 checks passed

This branch was successfully deployed

1 active deployment
Preview – cmux — c3e95f58 Deployed Jun 26, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

iOS: voice dictation mic button has animation lag (audio activation blocks main thread)

1 participant