Skip to content

Deliver cmux agent messages to Codex through its hooks - #15313

Merged
teamleaderleo merged 4 commits into
agent-inboxfrom
agent-inbox-codex
Sep 29, 2026
Merged

teamleaderleo merged 4 commits into
agent-inboxfrom
agent-inbox-codex

Conversation

@teamleaderleo

@teamleaderleo teamleaderleo commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Stacked on #15279 (review that first; this PR's diff is only the last commit). It gives Codex sessions the agent messages from cmux agent message. Part of manaflow-ai/cmuxterm-hq#847.

Codex hooks can't wake an idle session, so a message waits for Codex's next hook:

  • When the human submits a prompt, hooks codex inbox-drain attaches queued messages to it as additionalContext.
  • When Codex is about to stop, hooks codex inbox-stop returns {"decision":"block","reason":…}, so Codex continues the turn with the message instead of going idle.
  • An idle Codex sees the message at its next prompt. Nothing is ever typed into the terminal, so a draft in Codex's prompt is never touched.

How it's wired:

  • The UserPromptSubmit and Stop groups that the Codex wrapper injects now hold a second handler, run directly so its stdout reaches Codex. It is still one cmux group per event, as Codex wrapper injection replaces existing per-event hook arrays #12081 requires. The lifecycle handlers stay queued.
  • CodexHookInjectionEvent gets an optional companion, and configValue(command:) renders the -c value. Generation and the tests use that one renderer.
  • The previous schema moves into the exact recognized shapes, so saved launches from before this change still sanitize.
  • The replay sanitizer accepts the two-handler group only when both handlers are cmux's. If the second handler is someone else's, the whole block is kept.
  • codex exec runs get CMUX_CODEX_HEADLESS=1 from the wrapper, and the inbox handlers then return {}. So a headless run in an agent's pane can't take that agent's messages. claude -p works the same way in Agent messages that never land in a human's draft: cmux agent message #15279.

Not covered yet: sessions whose cmux hooks come from cmux hooks codex install instead of the launch wrapper. Their persistent events aren't re-injected, so they don't get the companions. The docs say so.

Testing

  • CodexHookInjectionStrippingTests and CodexHookPathSafetyTests: 30 tests pass on Swift 6.1.3 in a minimal scratch package with the sanitizer and schema sources. That run was on Linux, where the full CMUXAgentLaunch package can't build. New tests cover:
    • the companion shape;
    • stripping the pre-companion block;
    • keeping a block whose second handler is foreign.
  • CLICodexQueuedHookContractTests gains a check that hooks codex inject-args emits two handlers for UserPromptSubmit and Stop, the second direct, and one handler for the other events. It is not run locally (macOS only).
  • tests/test_codex_wrapper_hook_append.py now expects the companion in those two groups, and in its live codex exec run exactly one call to each inbox handler. CI does not provision Codex, so the live part is skipped there.
  • python3 scripts/verify-local.py --affected origin/main: 14/14 checks passed.
  • Not yet run against a live Codex in a tagged build; that will be part of the dogfood.

Changelog

Added: Codex sessions receive cmux agent message messages when the human next submits a prompt, or before Codex goes idle

Checklist

  • Behavior changes have added or updated tests, or Testing says why not
  • User-facing docs updated if needed
  • Reviewed with a subagent before merge, and all bot and human review comments resolved

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.


Summary by cubic

Delivers cmux agent message messages to Codex sessions through the injected Codex hooks. Since Codex hooks can't wake an idle session, a message waits for the next hook: it's attached as context to the human's next prompt, or Codex continues its turn with it instead of going idle.

  • The UserPromptSubmit and Stop hook groups carry a second direct handler for the inbox; the lifecycle handlers stay queued. The inbox handlers fail open, answering {} when the app is unreachable, so a dead app doesn't surface as a failed hook on every prompt.
  • Headless codex exec runs skip the inbox via a new CMUX_CODEX_HEADLESS env var, matching how claude -p behaves. The check accounts for every value-taking option, so codex --add-dir ../lib exec is treated as headless too.
  • The previous hook schema is preserved as a recognized shape, and the sanitizer accepts the two-handler groups only when both handlers are cmux's.
  • Sessions whose hooks come from cmux hooks codex install rather than the launch wrapper don't receive messages yet.

Written for commit 1e79f3c. Summary will update on new commits.

Review in cubic

@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository: manaflow-ai/cmux/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: ee83967a-9dd0-43b7-b695-72641d0bb935

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Dogfood build of 201d08f3c1a1cf8fc25986fe6540ce608a193c7e

cmux DEV pr-15313-201d08f3.app

The link opens this exact commit in the cmux dev menu bar app. The build starts on each push and the page waits until it is ready; a newer push replaces it. It signs in against production, so Cloud or backend changes still need a tagged build with a development backend.

@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Review: a review subagent read the diff and tested it against Codex 0.154 with a local fake model provider. Confirmed: one hook group with two command handlers is accepted, {} merges with additionalContext, and a Stop decision:block continues the turn with the message. The old recognized schema still strips pre-PR blocks, and 41 sanitizer tests pass, including probes that must not be stripped.

Fixed:

  • The headless check missed Codex options that take a value (--add-dir, -i/--image, --local-provider, --remote-auth-token-env), so codex --add-dir ../lib exec ... could take messages meant for the interactive Codex in the same pane. Regression test first (bf4afaa), then the fix. The same list decides hook injection, which is now more accurate too.
  • The message handlers did not fail open when the CLI itself failed (socket down, bad binary), so Codex showed "hook failed" on every prompt and stop while the app was unreachable. Regression test first (2cca0cf), then the fix.
  • Corrected the sanitizer comment that claimed the separator can only occur once.

Left:

  • Codex reads hook timeout in seconds, but cmux's injected -c config has always written milliseconds, so every injected timeout is 1000 times too long. This predates the PR. The persistent installer already converts. I'll fix it in a separate PR against main, since it changes every event's injected value and needs a new recognized schema.
  • The inline-command check in the sanitizer is a substring match, so a crafted user handler that mentions cmux's markers is dropped from replayed argv. This looseness already existed for single-handler values, and the only effect is that the user's handler isn't replayed.
  • Not verified: whether Codex fires Stop or UserPromptSubmit for subagent threads, which would hand them a message.
  • Not run here: the app and CLI targets and cmuxCLITests, which don't build on Linux. CI covers them.

@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Cross-model review (Codex gpt-5.6-sol)

  • CLI/CMUXCLI+CodexFireAndForgetHooks.swift:201-204 — wrapper generation suppresses an entire event when it finds any cmux-owned persistent hook for that event. Existing persistent installs add only the lifecycle handler for UserPromptSubmit/Stop, while this PR adds inbox-drain/inbox-stop only as wrapper companions; users who previously ran the persistent installer therefore lose the companion handler and queued agent messages are not delivered. Suppress per handler (inject a missing companion beside an existing lifecycle handler), or make persistent installation add both and suppress only when the complete pair is present; add a fixture for the current single-handler install.

@github-actions

Copy link
Copy Markdown
Contributor

CI failure attribution

CI failed on e0c8821c52 (run 36410774535 attempt 1): 1 machine.

Job Verdict Why
macos / app-host unit tests (changed suites) machine a binary loaded a framework from another build (stale products on the runner) (runner cmux10s-mac-mini-glaeda-4)
Matched log lines
macos / app-host unit tests (changed suites): ↳ status=6 reason=exit timedOut=false stdout=^D��dyld[24426]: Symbol not found: _$s15CMUXAgentLaunch05AgentB7CommandV10rejectedOn8launcher16externalLauncher14executablePath16workingDirectory11environment16verificationHome10capturedAt6sourceAcA0cB22CaptureRejectionReasonV_SSSgA3OSDyS2SGSgAOSdSgAOtcfC

Every failure is a machine failure: re-ran the failed jobs as attempt 2 (the checks show its result).

Written by scripts/ci/classify_failures.py (ci-failure-attribution.yml); signatures are its SIGNATURES table. A machine verdict is the runner's fault, not this PR's.

teamleaderleo and others added 4 commits September 29, 2026 10:01
The UserPromptSubmit and Stop hook groups the Codex wrapper injects now
carry a second, direct handler: inbox-drain attaches queued messages to
the prompt the human just sent, and inbox-stop continues the turn with
them instead of going idle. The lifecycle handlers stay queued.

The previous schema moves to the exact recognized shapes, and the replay
sanitizer accepts the two-handler group only when both handlers are
cmux's. codex exec runs skip the inbox, as claude -p runs do.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The headless check now knows every Codex option that takes a value, so
`codex --add-dir ../lib exec` no longer takes the pane's messages. The
agent message handlers answer {} when the CLI fails, so an unreachable
app does not show a failed hook on every prompt.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@teamleaderleo
teamleaderleo merged commit cae5aa9 into agent-inbox Sep 29, 2026
19 checks passed
@teamleaderleo
teamleaderleo deleted the agent-inbox-codex branch September 29, 2026 14:03
teamleaderleo added a commit that referenced this pull request Sep 30, 2026
…15863)

* Add a durable agent message store with hook-friendly waiting

Messages to the agent in a cmux surface are stored with queued, delivered
and read receipts in an append-only JSON Lines file, validated so no
control characters can ride along, and rendered once for every delivery
path with a header that marks the body as another agent's words. Waiters
are continuations, so a long-poll from a hook never parks a thread.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Deliver cmux agent messages through agent hooks, never keystrokes

cmux agent message <target> <text> stores a message for the agent in
another workspace or surface (agent.message.send). Claude Code gets it
through two new hooks instead of the terminal:

- hooks claude inbox-wait runs in the background (asyncRewake) after every
  session start and stop, long-polls agent.message.wait, and exits 2 with
  the message when one arrives. That wakes an idle session with the text
  as a system reminder and leaves the prompt box, and any half-typed
  draft in it, untouched.
- hooks claude inbox-drain on UserPromptSubmit attaches anything still
  queued as additionalContext when the human submits first. It fails open
  to {} and never exits 2, which would erase the prompt.

Delivery holds while the surface is waiting on a human (a question,
permission or plan prompt). agent.message.wait awaits a store
continuation on the socket worker, so a waiting hook never parks a
thread. Receipts (queued, delivered, read) go out on cmux events.

Codex handlers (inbox-drain, inbox-stop) are in place; wiring them into
the Codex launch schema is a follow-up. Remote relay stays denied.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Document cmux agent message and point agents at it

cmux docs agents, the agent help group and the cmux-workspace skill now
say to use cmux agent message instead of typing into another agent's
terminal. docs/agent-messages.md covers delivery, limits and the socket
API; docs/events.md lists the new receipts. Strings are localized for
all nine macOS locales.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Poll for agent messages instead of holding a socket connection

Addresses the review of the first pass:
- The Claude wake hook now checks agent.message.poll about every 2 seconds
  on a new connection, instead of a long poll that held one of the app's
  32 socket connection slots per session.
- The poll claims nothing; the hook claims right before handing messages
  to Claude, so a hook that died can no longer swallow them.
- The newest hook registers as the surface's poller; older ones (one per
  Stop) exit when superseded. Headless claude -p runs skip the inbox.
- The wait hook is also marked async, so a Claude Code without
  asyncRewake runs it in the background instead of blocking on it.
- A failed append never rewrites an existing message file.
- Options and -h after -- are message text; sender ids are canonical.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Replace namespace enums in the agent message package

The package conventions lint rejects all-static namespace types:
validation moves to AgentMessageDraft.validated() and rendering to
[AgentMessage].agentPromptText.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that a rendered agent message ends with its own id

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* End each rendered agent message with a line carrying its id

The chat view parses delivered messages out of agent transcripts. With a
bare --- as the end, a body quoting a message header could hide the real
message or fake its sender, and later hook output leaked into the body.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fail agent.message.send when the message can't be saved

AgentMessageStore.append now writes the journal record first and only
then adds the message to the in-memory inbox and fires onChange. An
encode, open, seek, write, or create failure throws
AgentMessagePersistenceError, and the socket command returns
storage_failed instead of reporting the message queued. A failed write
is truncated back off the file so a partial line can't swallow the next
record. State-change records stay best effort: losing one can only
repeat a delivery after restart.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: close Claude hook settings string after rebase

* Deliver cmux agent messages to Codex through its hooks (#15313)

* Deliver cmux agent messages to Codex through its hooks

The UserPromptSubmit and Stop hook groups the Codex wrapper injects now
carry a second, direct handler: inbox-drain attaches queued messages to
the prompt the human just sent, and inbox-stop continues the turn with
them instead of going idle. The lifecycle handlers stay queued.

The previous schema moves to the exact recognized shapes, and the replay
sanitizer accepts the two-handler group only when both handlers are
cmux's. codex exec runs skip the inbox, as claude -p runs do.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that value-taking Codex options do not hide codex exec

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that the Codex agent message handlers fail open

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Skip codex exec behind more options and fail open in message hooks

The headless check now knows every Codex option that takes a value, so
`codex --add-dir ../lib exec` no longer takes the pane's messages. The
agent message handlers answer {} when the CLI fails, so an unreachable
app does not show a failed hook on every prompt.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Show cmux agent messages in the terminal chat view (#15338)

* Test that the chat view shows delivered cmux agent messages

Fixtures use the transcript shapes Claude Code 2.1.283 and Codex 0.154
write for hook context, stop feedback, idle wakes and stop continuations.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that the chat view lists a terminal's queued agent messages

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Show cmux agent messages in the terminal chat view

Delivered messages appear as "Message from <sender>" in the turn they
arrived in, read from the agent's own transcript so they survive a
sidecar restart. Queued messages show above the composer, read from the
app every 2 seconds while a page is open.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test the chat view against quoted, forged and trailing message text

Covers the review findings: a body quoting a header, a forged message
inside a body, hook output after a message, a task result quoting a
message, a Codex prompt recorded twice, a failed queued read, the running
state after a wake, and a sender name with replacement patterns.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Parse agent messages by their id end line and tidy the queued view

Messages are read header by header and each ends at its own id line, so
quoted or forged text in a body and hook output after it stay out. Task
results are no longer parsed, a woken agent shows as running, Codex
prompts are not doubled by hook context, and a failed queued read keeps
the list and backs off. The queued panel scrolls past 30% of the view,
sender names are inserted literally, and the sessions list no longer
carries message bodies.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: drop removed Dock localization entry

* ci: rerun full app validation

* fix: handle agent message main actor hop failures

* fix: keep agent message timeout helpers local

* fix: address agent message review findings

* test: cover remote agent message relay policy

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover app-side agent message relay authorization

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: allow owned agent messages through SSH relay

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: track generated Claude hook groups

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: reject local splits from relay agent targets

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: close relay agent message authorization gaps

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover relayed agent message handlers

Add handler-level coverage for relay send and list authorization when a local split shares the workspace. Keep relay ID aliases synchronized as remote terminal mirror surfaces are added, removed, or untracked.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* dogfood: add agent message draft tour

Add a CI dogfood tour that shows an unsent terminal draft surviving agent message delivery and captures the queued message state.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* dogfood: show the queued inbox beside an unsent draft

Use the real CLI inbox in a separate pane. A fresh tour has no recorded terminal agent transcript, so opening its chat view cannot show a queued message. Keep all Return input confined to the new inbox pane and attach both surfaces' text snapshots.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: disambiguate agent message panel iteration

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: type agent message panel IDs

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Sync Ghostty pointer with main

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Sync Bonsplit pointer with main

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix agent journal test import

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix combined agent message and send guard merge

Preserve both localization key sets in a single union and keep the CLI command table parseable after combining the send guard and agent inbox documentation.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
teamleaderleo added a commit that referenced this pull request Sep 30, 2026
…#15279)

* Add a durable agent message store with hook-friendly waiting

Messages to the agent in a cmux surface are stored with queued, delivered
and read receipts in an append-only JSON Lines file, validated so no
control characters can ride along, and rendered once for every delivery
path with a header that marks the body as another agent's words. Waiters
are continuations, so a long-poll from a hook never parks a thread.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Deliver cmux agent messages through agent hooks, never keystrokes

cmux agent message <target> <text> stores a message for the agent in
another workspace or surface (agent.message.send). Claude Code gets it
through two new hooks instead of the terminal:

- hooks claude inbox-wait runs in the background (asyncRewake) after every
  session start and stop, long-polls agent.message.wait, and exits 2 with
  the message when one arrives. That wakes an idle session with the text
  as a system reminder and leaves the prompt box, and any half-typed
  draft in it, untouched.
- hooks claude inbox-drain on UserPromptSubmit attaches anything still
  queued as additionalContext when the human submits first. It fails open
  to {} and never exits 2, which would erase the prompt.

Delivery holds while the surface is waiting on a human (a question,
permission or plan prompt). agent.message.wait awaits a store
continuation on the socket worker, so a waiting hook never parks a
thread. Receipts (queued, delivered, read) go out on cmux events.

Codex handlers (inbox-drain, inbox-stop) are in place; wiring them into
the Codex launch schema is a follow-up. Remote relay stays denied.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Document cmux agent message and point agents at it

cmux docs agents, the agent help group and the cmux-workspace skill now
say to use cmux agent message instead of typing into another agent's
terminal. docs/agent-messages.md covers delivery, limits and the socket
API; docs/events.md lists the new receipts. Strings are localized for
all nine macOS locales.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Poll for agent messages instead of holding a socket connection

Addresses the review of the first pass:
- The Claude wake hook now checks agent.message.poll about every 2 seconds
  on a new connection, instead of a long poll that held one of the app's
  32 socket connection slots per session.
- The poll claims nothing; the hook claims right before handing messages
  to Claude, so a hook that died can no longer swallow them.
- The newest hook registers as the surface's poller; older ones (one per
  Stop) exit when superseded. Headless claude -p runs skip the inbox.
- The wait hook is also marked async, so a Claude Code without
  asyncRewake runs it in the background instead of blocking on it.
- A failed append never rewrites an existing message file.
- Options and -h after -- are message text; sender ids are canonical.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Replace namespace enums in the agent message package

The package conventions lint rejects all-static namespace types:
validation moves to AgentMessageDraft.validated() and rendering to
[AgentMessage].agentPromptText.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that a rendered agent message ends with its own id

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* End each rendered agent message with a line carrying its id

The chat view parses delivered messages out of agent transcripts. With a
bare --- as the end, a body quoting a message header could hide the real
message or fake its sender, and later hook output leaked into the body.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fail agent.message.send when the message can't be saved

AgentMessageStore.append now writes the journal record first and only
then adds the message to the in-memory inbox and fires onChange. An
encode, open, seek, write, or create failure throws
AgentMessagePersistenceError, and the socket command returns
storage_failed instead of reporting the message queued. A failed write
is truncated back off the file so a partial line can't swallow the next
record. State-change records stay best effort: losing one can only
repeat a delivery after restart.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: close Claude hook settings string after rebase

* Deliver cmux agent messages to Codex through its hooks (#15313)

* Deliver cmux agent messages to Codex through its hooks

The UserPromptSubmit and Stop hook groups the Codex wrapper injects now
carry a second, direct handler: inbox-drain attaches queued messages to
the prompt the human just sent, and inbox-stop continues the turn with
them instead of going idle. The lifecycle handlers stay queued.

The previous schema moves to the exact recognized shapes, and the replay
sanitizer accepts the two-handler group only when both handlers are
cmux's. codex exec runs skip the inbox, as claude -p runs do.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that value-taking Codex options do not hide codex exec

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that the Codex agent message handlers fail open

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Skip codex exec behind more options and fail open in message hooks

The headless check now knows every Codex option that takes a value, so
`codex --add-dir ../lib exec` no longer takes the pane's messages. The
agent message handlers answer {} when the CLI fails, so an unreachable
app does not show a failed hook on every prompt.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Show cmux agent messages in the terminal chat view (#15338)

* Test that the chat view shows delivered cmux agent messages

Fixtures use the transcript shapes Claude Code 2.1.283 and Codex 0.154
write for hook context, stop feedback, idle wakes and stop continuations.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that the chat view lists a terminal's queued agent messages

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Show cmux agent messages in the terminal chat view

Delivered messages appear as "Message from <sender>" in the turn they
arrived in, read from the agent's own transcript so they survive a
sidecar restart. Queued messages show above the composer, read from the
app every 2 seconds while a page is open.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test the chat view against quoted, forged and trailing message text

Covers the review findings: a body quoting a header, a forged message
inside a body, hook output after a message, a task result quoting a
message, a Codex prompt recorded twice, a failed queued read, the running
state after a wake, and a sender name with replacement patterns.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Parse agent messages by their id end line and tidy the queued view

Messages are read header by header and each ends at its own id line, so
quoted or forged text in a body and hook output after it stay out. Task
results are no longer parsed, a woken agent shows as running, Codex
prompts are not doubled by hook context, and a failed queued read keeps
the list and backs off. The queued panel scrolls past 30% of the view,
sender names are inserted literally, and the sessions list no longer
carries message bodies.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: drop removed Dock localization entry

* ci: rerun full app validation

* fix: handle agent message main actor hop failures

* fix: keep agent message timeout helpers local

* fix: address agent message review findings

* test: track generated Claude hook groups

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* dogfood: add agent message draft tour

Add a CI dogfood tour that shows an unsent terminal draft surviving agent message delivery and captures the queued message state.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* dogfood: show the queued inbox beside an unsent draft

Use the real CLI inbox in a separate pane. A fresh tour has no recorded terminal agent transcript, so opening its chat view cannot show a queued message. Keep all Return input confined to the new inbox pane and attach both surfaces' text snapshots.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent message review regressions

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: address agent message review findings

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(agent-chat): place shared message type at package root

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent message poller ownership

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: bind deferred agent delivery to its poller

Deferred wake reads now filter by recipient and verify the active poller. The rendered sender header matches the chat parser so delivered messages show the sender name directly.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover StopFailure inbox wake hook

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover deferred agent message reservations

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(agent-messages): make deferred wake delivery retryable

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
teamleaderleo added a commit that referenced this pull request Oct 1, 2026
* Add a durable agent message store with hook-friendly waiting

Messages to the agent in a cmux surface are stored with queued, delivered
and read receipts in an append-only JSON Lines file, validated so no
control characters can ride along, and rendered once for every delivery
path with a header that marks the body as another agent's words. Waiters
are continuations, so a long-poll from a hook never parks a thread.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Deliver cmux agent messages through agent hooks, never keystrokes

cmux agent message <target> <text> stores a message for the agent in
another workspace or surface (agent.message.send). Claude Code gets it
through two new hooks instead of the terminal:

- hooks claude inbox-wait runs in the background (asyncRewake) after every
  session start and stop, long-polls agent.message.wait, and exits 2 with
  the message when one arrives. That wakes an idle session with the text
  as a system reminder and leaves the prompt box, and any half-typed
  draft in it, untouched.
- hooks claude inbox-drain on UserPromptSubmit attaches anything still
  queued as additionalContext when the human submits first. It fails open
  to {} and never exits 2, which would erase the prompt.

Delivery holds while the surface is waiting on a human (a question,
permission or plan prompt). agent.message.wait awaits a store
continuation on the socket worker, so a waiting hook never parks a
thread. Receipts (queued, delivered, read) go out on cmux events.

Codex handlers (inbox-drain, inbox-stop) are in place; wiring them into
the Codex launch schema is a follow-up. Remote relay stays denied.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Document cmux agent message and point agents at it

cmux docs agents, the agent help group and the cmux-workspace skill now
say to use cmux agent message instead of typing into another agent's
terminal. docs/agent-messages.md covers delivery, limits and the socket
API; docs/events.md lists the new receipts. Strings are localized for
all nine macOS locales.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Poll for agent messages instead of holding a socket connection

Addresses the review of the first pass:
- The Claude wake hook now checks agent.message.poll about every 2 seconds
  on a new connection, instead of a long poll that held one of the app's
  32 socket connection slots per session.
- The poll claims nothing; the hook claims right before handing messages
  to Claude, so a hook that died can no longer swallow them.
- The newest hook registers as the surface's poller; older ones (one per
  Stop) exit when superseded. Headless claude -p runs skip the inbox.
- The wait hook is also marked async, so a Claude Code without
  asyncRewake runs it in the background instead of blocking on it.
- A failed append never rewrites an existing message file.
- Options and -h after -- are message text; sender ids are canonical.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Replace namespace enums in the agent message package

The package conventions lint rejects all-static namespace types:
validation moves to AgentMessageDraft.validated() and rendering to
[AgentMessage].agentPromptText.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that a rendered agent message ends with its own id

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* End each rendered agent message with a line carrying its id

The chat view parses delivered messages out of agent transcripts. With a
bare --- as the end, a body quoting a message header could hide the real
message or fake its sender, and later hook output leaked into the body.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fail agent.message.send when the message can't be saved

AgentMessageStore.append now writes the journal record first and only
then adds the message to the in-memory inbox and fires onChange. An
encode, open, seek, write, or create failure throws
AgentMessagePersistenceError, and the socket command returns
storage_failed instead of reporting the message queued. A failed write
is truncated back off the file so a partial line can't swallow the next
record. State-change records stay best effort: losing one can only
repeat a delivery after restart.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: close Claude hook settings string after rebase

* Deliver cmux agent messages to Codex through its hooks (#15313)

* Deliver cmux agent messages to Codex through its hooks

The UserPromptSubmit and Stop hook groups the Codex wrapper injects now
carry a second, direct handler: inbox-drain attaches queued messages to
the prompt the human just sent, and inbox-stop continues the turn with
them instead of going idle. The lifecycle handlers stay queued.

The previous schema moves to the exact recognized shapes, and the replay
sanitizer accepts the two-handler group only when both handlers are
cmux's. codex exec runs skip the inbox, as claude -p runs do.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that value-taking Codex options do not hide codex exec

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that the Codex agent message handlers fail open

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Skip codex exec behind more options and fail open in message hooks

The headless check now knows every Codex option that takes a value, so
`codex --add-dir ../lib exec` no longer takes the pane's messages. The
agent message handlers answer {} when the CLI fails, so an unreachable
app does not show a failed hook on every prompt.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Show cmux agent messages in the terminal chat view (#15338)

* Test that the chat view shows delivered cmux agent messages

Fixtures use the transcript shapes Claude Code 2.1.283 and Codex 0.154
write for hook context, stop feedback, idle wakes and stop continuations.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that the chat view lists a terminal's queued agent messages

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Show cmux agent messages in the terminal chat view

Delivered messages appear as "Message from <sender>" in the turn they
arrived in, read from the agent's own transcript so they survive a
sidecar restart. Queued messages show above the composer, read from the
app every 2 seconds while a page is open.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test the chat view against quoted, forged and trailing message text

Covers the review findings: a body quoting a header, a forged message
inside a body, hook output after a message, a task result quoting a
message, a Codex prompt recorded twice, a failed queued read, the running
state after a wake, and a sender name with replacement patterns.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Parse agent messages by their id end line and tidy the queued view

Messages are read header by header and each ends at its own id line, so
quoted or forged text in a body and hook output after it stay out. Task
results are no longer parsed, a woken agent shows as running, Codex
prompts are not doubled by hook context, and a failed queued read keeps
the list and backs off. The queued panel scrolls past 30% of the view,
sender names are inserted literally, and the sessions list no longer
carries message bodies.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: drop removed Dock localization entry

* ci: rerun full app validation

* fix: handle agent message main actor hop failures

* fix: keep agent message timeout helpers local

* fix: address agent message review findings

* test: define agent inbox projection behavior

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: track generated Claude hook groups

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: add agent inbox quick view

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent inbox crash and reply path

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: harden agent inbox quick view

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent inbox review regressions

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: harden agent inbox quick view interactions

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep agent inbox shortcut keymaps conflict-free

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover remaining agent inbox review items

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: address agent inbox review items

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: add Agent Inbox dogfood tour

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent inbox focus notification window

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: report the agent inbox hosting window

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: compile reply gate assertions with Swift Testing

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent inbox review regressions first

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: address agent inbox review feedback

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* chore: advance submodules with main

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix main merge artifacts

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: pass auto-naming config mode to provider overrides

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep OpenCode path resolution in the CLI target

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
austinywang pushed a commit that referenced this pull request Oct 1, 2026
* Add a durable agent message store with hook-friendly waiting

Messages to the agent in a cmux surface are stored with queued, delivered
and read receipts in an append-only JSON Lines file, validated so no
control characters can ride along, and rendered once for every delivery
path with a header that marks the body as another agent's words. Waiters
are continuations, so a long-poll from a hook never parks a thread.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Deliver cmux agent messages through agent hooks, never keystrokes

cmux agent message <target> <text> stores a message for the agent in
another workspace or surface (agent.message.send). Claude Code gets it
through two new hooks instead of the terminal:

- hooks claude inbox-wait runs in the background (asyncRewake) after every
  session start and stop, long-polls agent.message.wait, and exits 2 with
  the message when one arrives. That wakes an idle session with the text
  as a system reminder and leaves the prompt box, and any half-typed
  draft in it, untouched.
- hooks claude inbox-drain on UserPromptSubmit attaches anything still
  queued as additionalContext when the human submits first. It fails open
  to {} and never exits 2, which would erase the prompt.

Delivery holds while the surface is waiting on a human (a question,
permission or plan prompt). agent.message.wait awaits a store
continuation on the socket worker, so a waiting hook never parks a
thread. Receipts (queued, delivered, read) go out on cmux events.

Codex handlers (inbox-drain, inbox-stop) are in place; wiring them into
the Codex launch schema is a follow-up. Remote relay stays denied.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Document cmux agent message and point agents at it

cmux docs agents, the agent help group and the cmux-workspace skill now
say to use cmux agent message instead of typing into another agent's
terminal. docs/agent-messages.md covers delivery, limits and the socket
API; docs/events.md lists the new receipts. Strings are localized for
all nine macOS locales.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Poll for agent messages instead of holding a socket connection

Addresses the review of the first pass:
- The Claude wake hook now checks agent.message.poll about every 2 seconds
  on a new connection, instead of a long poll that held one of the app's
  32 socket connection slots per session.
- The poll claims nothing; the hook claims right before handing messages
  to Claude, so a hook that died can no longer swallow them.
- The newest hook registers as the surface's poller; older ones (one per
  Stop) exit when superseded. Headless claude -p runs skip the inbox.
- The wait hook is also marked async, so a Claude Code without
  asyncRewake runs it in the background instead of blocking on it.
- A failed append never rewrites an existing message file.
- Options and -h after -- are message text; sender ids are canonical.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Replace namespace enums in the agent message package

The package conventions lint rejects all-static namespace types:
validation moves to AgentMessageDraft.validated() and rendering to
[AgentMessage].agentPromptText.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that a rendered agent message ends with its own id

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* End each rendered agent message with a line carrying its id

The chat view parses delivered messages out of agent transcripts. With a
bare --- as the end, a body quoting a message header could hide the real
message or fake its sender, and later hook output leaked into the body.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fail agent.message.send when the message can't be saved

AgentMessageStore.append now writes the journal record first and only
then adds the message to the in-memory inbox and fires onChange. An
encode, open, seek, write, or create failure throws
AgentMessagePersistenceError, and the socket command returns
storage_failed instead of reporting the message queued. A failed write
is truncated back off the file so a partial line can't swallow the next
record. State-change records stay best effort: losing one can only
repeat a delivery after restart.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: close Claude hook settings string after rebase

* Deliver cmux agent messages to Codex through its hooks (#15313)

* Deliver cmux agent messages to Codex through its hooks

The UserPromptSubmit and Stop hook groups the Codex wrapper injects now
carry a second, direct handler: inbox-drain attaches queued messages to
the prompt the human just sent, and inbox-stop continues the turn with
them instead of going idle. The lifecycle handlers stay queued.

The previous schema moves to the exact recognized shapes, and the replay
sanitizer accepts the two-handler group only when both handlers are
cmux's. codex exec runs skip the inbox, as claude -p runs do.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that value-taking Codex options do not hide codex exec

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that the Codex agent message handlers fail open

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Skip codex exec behind more options and fail open in message hooks

The headless check now knows every Codex option that takes a value, so
`codex --add-dir ../lib exec` no longer takes the pane's messages. The
agent message handlers answer {} when the CLI fails, so an unreachable
app does not show a failed hook on every prompt.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Show cmux agent messages in the terminal chat view (#15338)

* Test that the chat view shows delivered cmux agent messages

Fixtures use the transcript shapes Claude Code 2.1.283 and Codex 0.154
write for hook context, stop feedback, idle wakes and stop continuations.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that the chat view lists a terminal's queued agent messages

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Show cmux agent messages in the terminal chat view

Delivered messages appear as "Message from <sender>" in the turn they
arrived in, read from the agent's own transcript so they survive a
sidecar restart. Queued messages show above the composer, read from the
app every 2 seconds while a page is open.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test the chat view against quoted, forged and trailing message text

Covers the review findings: a body quoting a header, a forged message
inside a body, hook output after a message, a task result quoting a
message, a Codex prompt recorded twice, a failed queued read, the running
state after a wake, and a sender name with replacement patterns.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Parse agent messages by their id end line and tidy the queued view

Messages are read header by header and each ends at its own id line, so
quoted or forged text in a body and hook output after it stay out. Task
results are no longer parsed, a woken agent shows as running, Codex
prompts are not doubled by hook context, and a failed queued read keeps
the list and backs off. The queued panel scrolls past 30% of the view,
sender names are inserted literally, and the sessions list no longer
carries message bodies.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: drop removed Dock localization entry

* ci: rerun full app validation

* fix: handle agent message main actor hop failures

* fix: keep agent message timeout helpers local

* fix: address agent message review findings

* test: define agent inbox projection behavior

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: track generated Claude hook groups

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: add agent inbox quick view

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent inbox crash and reply path

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: harden agent inbox quick view

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent inbox review regressions

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: harden agent inbox quick view interactions

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep agent inbox shortcut keymaps conflict-free

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover remaining agent inbox review items

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: address agent inbox review items

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: add Agent Inbox dogfood tour

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent inbox focus notification window

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: report the agent inbox hosting window

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: compile reply gate assertions with Swift Testing

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent inbox review regressions first

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: address agent inbox review feedback

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* chore: advance submodules with main

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix main merge artifacts

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: pass auto-naming config mode to provider overrides

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep OpenCode path resolution in the CLI target

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
austinywang pushed a commit that referenced this pull request Oct 1, 2026
* Add a durable agent message store with hook-friendly waiting

Messages to the agent in a cmux surface are stored with queued, delivered
and read receipts in an append-only JSON Lines file, validated so no
control characters can ride along, and rendered once for every delivery
path with a header that marks the body as another agent's words. Waiters
are continuations, so a long-poll from a hook never parks a thread.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Deliver cmux agent messages through agent hooks, never keystrokes

cmux agent message <target> <text> stores a message for the agent in
another workspace or surface (agent.message.send). Claude Code gets it
through two new hooks instead of the terminal:

- hooks claude inbox-wait runs in the background (asyncRewake) after every
  session start and stop, long-polls agent.message.wait, and exits 2 with
  the message when one arrives. That wakes an idle session with the text
  as a system reminder and leaves the prompt box, and any half-typed
  draft in it, untouched.
- hooks claude inbox-drain on UserPromptSubmit attaches anything still
  queued as additionalContext when the human submits first. It fails open
  to {} and never exits 2, which would erase the prompt.

Delivery holds while the surface is waiting on a human (a question,
permission or plan prompt). agent.message.wait awaits a store
continuation on the socket worker, so a waiting hook never parks a
thread. Receipts (queued, delivered, read) go out on cmux events.

Codex handlers (inbox-drain, inbox-stop) are in place; wiring them into
the Codex launch schema is a follow-up. Remote relay stays denied.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Document cmux agent message and point agents at it

cmux docs agents, the agent help group and the cmux-workspace skill now
say to use cmux agent message instead of typing into another agent's
terminal. docs/agent-messages.md covers delivery, limits and the socket
API; docs/events.md lists the new receipts. Strings are localized for
all nine macOS locales.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Poll for agent messages instead of holding a socket connection

Addresses the review of the first pass:
- The Claude wake hook now checks agent.message.poll about every 2 seconds
  on a new connection, instead of a long poll that held one of the app's
  32 socket connection slots per session.
- The poll claims nothing; the hook claims right before handing messages
  to Claude, so a hook that died can no longer swallow them.
- The newest hook registers as the surface's poller; older ones (one per
  Stop) exit when superseded. Headless claude -p runs skip the inbox.
- The wait hook is also marked async, so a Claude Code without
  asyncRewake runs it in the background instead of blocking on it.
- A failed append never rewrites an existing message file.
- Options and -h after -- are message text; sender ids are canonical.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Replace namespace enums in the agent message package

The package conventions lint rejects all-static namespace types:
validation moves to AgentMessageDraft.validated() and rendering to
[AgentMessage].agentPromptText.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that a rendered agent message ends with its own id

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* End each rendered agent message with a line carrying its id

The chat view parses delivered messages out of agent transcripts. With a
bare --- as the end, a body quoting a message header could hide the real
message or fake its sender, and later hook output leaked into the body.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fail agent.message.send when the message can't be saved

AgentMessageStore.append now writes the journal record first and only
then adds the message to the in-memory inbox and fires onChange. An
encode, open, seek, write, or create failure throws
AgentMessagePersistenceError, and the socket command returns
storage_failed instead of reporting the message queued. A failed write
is truncated back off the file so a partial line can't swallow the next
record. State-change records stay best effort: losing one can only
repeat a delivery after restart.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: close Claude hook settings string after rebase

* Deliver cmux agent messages to Codex through its hooks (#15313)

* Deliver cmux agent messages to Codex through its hooks

The UserPromptSubmit and Stop hook groups the Codex wrapper injects now
carry a second, direct handler: inbox-drain attaches queued messages to
the prompt the human just sent, and inbox-stop continues the turn with
them instead of going idle. The lifecycle handlers stay queued.

The previous schema moves to the exact recognized shapes, and the replay
sanitizer accepts the two-handler group only when both handlers are
cmux's. codex exec runs skip the inbox, as claude -p runs do.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that value-taking Codex options do not hide codex exec

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that the Codex agent message handlers fail open

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Skip codex exec behind more options and fail open in message hooks

The headless check now knows every Codex option that takes a value, so
`codex --add-dir ../lib exec` no longer takes the pane's messages. The
agent message handlers answer {} when the CLI fails, so an unreachable
app does not show a failed hook on every prompt.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Show cmux agent messages in the terminal chat view (#15338)

* Test that the chat view shows delivered cmux agent messages

Fixtures use the transcript shapes Claude Code 2.1.283 and Codex 0.154
write for hook context, stop feedback, idle wakes and stop continuations.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that the chat view lists a terminal's queued agent messages

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Show cmux agent messages in the terminal chat view

Delivered messages appear as "Message from <sender>" in the turn they
arrived in, read from the agent's own transcript so they survive a
sidecar restart. Queued messages show above the composer, read from the
app every 2 seconds while a page is open.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test the chat view against quoted, forged and trailing message text

Covers the review findings: a body quoting a header, a forged message
inside a body, hook output after a message, a task result quoting a
message, a Codex prompt recorded twice, a failed queued read, the running
state after a wake, and a sender name with replacement patterns.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Parse agent messages by their id end line and tidy the queued view

Messages are read header by header and each ends at its own id line, so
quoted or forged text in a body and hook output after it stay out. Task
results are no longer parsed, a woken agent shows as running, Codex
prompts are not doubled by hook context, and a failed queued read keeps
the list and backs off. The queued panel scrolls past 30% of the view,
sender names are inserted literally, and the sessions list no longer
carries message bodies.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: drop removed Dock localization entry

* ci: rerun full app validation

* fix: handle agent message main actor hop failures

* fix: keep agent message timeout helpers local

* fix: address agent message review findings

* test: define agent inbox projection behavior

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: track generated Claude hook groups

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: add agent inbox quick view

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent inbox crash and reply path

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: harden agent inbox quick view

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent inbox review regressions

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: harden agent inbox quick view interactions

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep agent inbox shortcut keymaps conflict-free

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover remaining agent inbox review items

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: address agent inbox review items

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: add Agent Inbox dogfood tour

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent inbox focus notification window

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: report the agent inbox hosting window

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: compile reply gate assertions with Swift Testing

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent inbox review regressions first

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: address agent inbox review feedback

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* chore: advance submodules with main

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix main merge artifacts

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: pass auto-naming config mode to provider overrides

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep OpenCode path resolution in the CLI target

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
austinywang added a commit that referenced this pull request Oct 1, 2026
* fix: keep cloud terminals alive after journal failure

* ci: archive docs uploads for Vercel file limit (#16289)

* Add an agent inbox quick view behind a feature flag (#15888)

* Add a durable agent message store with hook-friendly waiting

Messages to the agent in a cmux surface are stored with queued, delivered
and read receipts in an append-only JSON Lines file, validated so no
control characters can ride along, and rendered once for every delivery
path with a header that marks the body as another agent's words. Waiters
are continuations, so a long-poll from a hook never parks a thread.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Deliver cmux agent messages through agent hooks, never keystrokes

cmux agent message <target> <text> stores a message for the agent in
another workspace or surface (agent.message.send). Claude Code gets it
through two new hooks instead of the terminal:

- hooks claude inbox-wait runs in the background (asyncRewake) after every
  session start and stop, long-polls agent.message.wait, and exits 2 with
  the message when one arrives. That wakes an idle session with the text
  as a system reminder and leaves the prompt box, and any half-typed
  draft in it, untouched.
- hooks claude inbox-drain on UserPromptSubmit attaches anything still
  queued as additionalContext when the human submits first. It fails open
  to {} and never exits 2, which would erase the prompt.

Delivery holds while the surface is waiting on a human (a question,
permission or plan prompt). agent.message.wait awaits a store
continuation on the socket worker, so a waiting hook never parks a
thread. Receipts (queued, delivered, read) go out on cmux events.

Codex handlers (inbox-drain, inbox-stop) are in place; wiring them into
the Codex launch schema is a follow-up. Remote relay stays denied.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Document cmux agent message and point agents at it

cmux docs agents, the agent help group and the cmux-workspace skill now
say to use cmux agent message instead of typing into another agent's
terminal. docs/agent-messages.md covers delivery, limits and the socket
API; docs/events.md lists the new receipts. Strings are localized for
all nine macOS locales.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Poll for agent messages instead of holding a socket connection

Addresses the review of the first pass:
- The Claude wake hook now checks agent.message.poll about every 2 seconds
  on a new connection, instead of a long poll that held one of the app's
  32 socket connection slots per session.
- The poll claims nothing; the hook claims right before handing messages
  to Claude, so a hook that died can no longer swallow them.
- The newest hook registers as the surface's poller; older ones (one per
  Stop) exit when superseded. Headless claude -p runs skip the inbox.
- The wait hook is also marked async, so a Claude Code without
  asyncRewake runs it in the background instead of blocking on it.
- A failed append never rewrites an existing message file.
- Options and -h after -- are message text; sender ids are canonical.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Replace namespace enums in the agent message package

The package conventions lint rejects all-static namespace types:
validation moves to AgentMessageDraft.validated() and rendering to
[AgentMessage].agentPromptText.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that a rendered agent message ends with its own id

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* End each rendered agent message with a line carrying its id

The chat view parses delivered messages out of agent transcripts. With a
bare --- as the end, a body quoting a message header could hide the real
message or fake its sender, and later hook output leaked into the body.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fail agent.message.send when the message can't be saved

AgentMessageStore.append now writes the journal record first and only
then adds the message to the in-memory inbox and fires onChange. An
encode, open, seek, write, or create failure throws
AgentMessagePersistenceError, and the socket command returns
storage_failed instead of reporting the message queued. A failed write
is truncated back off the file so a partial line can't swallow the next
record. State-change records stay best effort: losing one can only
repeat a delivery after restart.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: close Claude hook settings string after rebase

* Deliver cmux agent messages to Codex through its hooks (#15313)

* Deliver cmux agent messages to Codex through its hooks

The UserPromptSubmit and Stop hook groups the Codex wrapper injects now
carry a second, direct handler: inbox-drain attaches queued messages to
the prompt the human just sent, and inbox-stop continues the turn with
them instead of going idle. The lifecycle handlers stay queued.

The previous schema moves to the exact recognized shapes, and the replay
sanitizer accepts the two-handler group only when both handlers are
cmux's. codex exec runs skip the inbox, as claude -p runs do.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that value-taking Codex options do not hide codex exec

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that the Codex agent message handlers fail open

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Skip codex exec behind more options and fail open in message hooks

The headless check now knows every Codex option that takes a value, so
`codex --add-dir ../lib exec` no longer takes the pane's messages. The
agent message handlers answer {} when the CLI fails, so an unreachable
app does not show a failed hook on every prompt.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Show cmux agent messages in the terminal chat view (#15338)

* Test that the chat view shows delivered cmux agent messages

Fixtures use the transcript shapes Claude Code 2.1.283 and Codex 0.154
write for hook context, stop feedback, idle wakes and stop continuations.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test that the chat view lists a terminal's queued agent messages

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Show cmux agent messages in the terminal chat view

Delivered messages appear as "Message from <sender>" in the turn they
arrived in, read from the agent's own transcript so they survive a
sidecar restart. Queued messages show above the composer, read from the
app every 2 seconds while a page is open.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test the chat view against quoted, forged and trailing message text

Covers the review findings: a body quoting a header, a forged message
inside a body, hook output after a message, a task result quoting a
message, a Codex prompt recorded twice, a failed queued read, the running
state after a wake, and a sender name with replacement patterns.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Parse agent messages by their id end line and tidy the queued view

Messages are read header by header and each ends at its own id line, so
quoted or forged text in a body and hook output after it stay out. Task
results are no longer parsed, a woken agent shows as running, Codex
prompts are not doubled by hook context, and a failed queued read keeps
the list and backs off. The queued panel scrolls past 30% of the view,
sender names are inserted literally, and the sessions list no longer
carries message bodies.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: drop removed Dock localization entry

* ci: rerun full app validation

* fix: handle agent message main actor hop failures

* fix: keep agent message timeout helpers local

* fix: address agent message review findings

* test: define agent inbox projection behavior

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: track generated Claude hook groups

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: add agent inbox quick view

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent inbox crash and reply path

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: harden agent inbox quick view

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent inbox review regressions

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: harden agent inbox quick view interactions

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep agent inbox shortcut keymaps conflict-free

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover remaining agent inbox review items

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: address agent inbox review items

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: add Agent Inbox dogfood tour

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent inbox focus notification window

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: report the agent inbox hosting window

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: compile reply gate assertions with Swift Testing

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: cover agent inbox review regressions first

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: address agent inbox review feedback

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* chore: advance submodules with main

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix main merge artifacts

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: pass auto-naming config mode to provider overrides

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep OpenCode path resolution in the CLI target

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Update computer use engine for unrestricted app access

* fix: redact terminal output read failures

* test: preserve explicit shutdown assertion

* fix: identify terminal output read failures

---------

Co-authored-by: Austin Wang <austinwang115@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Lawrence Chen <54008264+lawrencecchen@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant