Skip to content

Device-list authorization + DO-first control plane (listauth) - #11220

Merged
azooz2003-bit merged 57 commits into
mainfrom
feat-listauth
Sep 1, 2026
Merged

azooz2003-bit merged 57 commits into
mainfrom
feat-listauth

Conversation

@azooz2003-bit

@azooz2003-bit azooz2003-bit commented Aug 30, 2026 •

Copy link
Copy Markdown
Collaborator

Implements the connection-architecture redesign specified in the living design artifact (hq cmux-assets/irx-client-resilience/irx-sequence-diagram): device-list authorization replaces pair-grant JWS verification, and the account Durable Object becomes the control plane for registration confirmation, device-list sync, acks, and revocation.

Backend (workers/presence): listv2 overlay in the AccountControlPlane DO (status enum + revoked + appVersion/track/capabilities per device, seeded-on-first-sight, confirm-on-hello), directory freshness lease (issuedAt + 24h ttl, snapshot_complete re-stamps), per-device ack(rev) with alarm retries (5s/30s/2m/10m/hourly), POST /v1/control/devices/revoke (instant broadcast, socket close 1008, mint denial). Schemas extended additively + quicktype regenerated (TS + Swift), drift guard green. 240 worker tests.

Clients (both platforms): grantless hello v2; Mac admission = IrxListJudge (EndpointID in list AND !revoked AND lease fresh, synchronous O(1), fail closed on stale/absent); revocation closes live sessions; iOS dial gate + seeded-Mac warning row (en+ja); device-list lease persisted in Keychain (ThisDeviceOnly, keyed account|backend); broker JSON caches move to Keychain in Release with one-way file migration; control-plane client sends client info + acks and starts before broker calls. 62 + 338 package tests.

Transport hardening found by the gates: keepalive now requires two consecutive pong misses before declaring death (a lone 2s stall severed an otherwise-perfect 9-min session; the Mac's pong landed 85ms past the deadline), and automatic path mode now actually authorizes NAT traversal after admission (deferNatTraversalUntilAuthorized was set at bind but nothing ever authorized, so sessions could never leave the relay — the known "direct-path unverified" v1 gap).

Also fixes a main compile break: AgentHibernationRecord.processLiveness had an inline default, excluding it from the memberwise init used by +Records.swift (#10658 fallout; the PR lane has no macOS compile gate).

Acceptance gates run live (isolated sim as iPhone, tagged Mac, isolated dev worker cmux-presence-dev-listauth proxying staging; harness scripts/listauth-gates.py; evidence in hq cmux-assets/irx-client-resilience/listauth-gates-evidence/):

  • 30-min relay soak PASS: 338/338 relay pongs, 0 reconnects/disconnects, 77/70 pass rotations and 4 list pushes crossed with zero session impact, 60 inputs, establish 93ms.
  • 30-min direct/LAN soak PASS: 344/344 pongs on the direct path (rtt 0ms after NAT-traversal upgrade), 0 breaks, 77/77 rotations, establish 70ms.
  • Cold start PASS: directory ("see all computers") 711–1313ms (<2s), session admitted 1457–2105ms (<3s transport target). Timing anchored at the signed-in moment because the sim re-runs a ~15s UITEST password sign-in every launch (dev-env artifact; product devices restore sessions locally).
  • Background-return gate (30-min hold) running at PR time; result posted as a comment.

Deferred by design decision: the full sign-in→terminal XCUITest (post-implementation design pass), hint-rev acks, Mac account-erase seam for broker caches (pre-existing), MobileMacListAuthState singleton seam.

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.


Summary by cubic

Replaces pair-grant JWS verification with device-list authorization and moves directory sync, acks, and revocation into the per-account control-plane Durable Object. Admission checks the Mac's endpoint against a persisted 24h-leased device list; revoked devices are cut immediately and denied future admission, and clients ack applied control-plane revisions so the server's retry ladder stops at them.

Transport and client behavior

  • Keepalive now requires two consecutive pong misses; automatic path mode authorizes NAT traversal after admission; foreground resume replaces zombie sessions immediately; explicit redial invalidates the old session first; the control lane no longer closes the shared QUIC session; reconnect liveness and lease freshness run on the monotonic clock so wall-clock rollback can't skip a zombie or revive a stale lease, and the persisted lease TTL is capped so a corrupt snapshot can't overflow the refresh math.
  • The broker mint retries once on a stale kept-alive connection and keeps the URL-error cause; Release broker caches move to the Keychain behind a one-way, account- and backend-scoped migration that preserves data and rejects replayed or wall-rolled-back leases; Settings Networking now reflects the irx runtime and throws on unsupported mutations.
  • Pairing QR codes carry the Mac's device ID, pairing attempts clear timed-out auth/token-suppression windows so fresh pairings persist their routes, and dev Macs publish registry routes to shared staging.

Control plane and infrastructure

  • Control-plane wire types are generated from schemas/control-plane/ with a CI drift guard; transport state is namespaced per bundle and broker so dev and production builds can't overwrite each other's trust keys.
  • The socket sends the app namespace so the DO presents devices hidden by dev-tag isolation; serialized control loops share no decoder state, recover from zombie sockets and snapshot-registry closes, and merge relay updates without double-minting; the dev-launch attach timeout rises to 60s for cold auth.
  • Exposes AgentHibernationRecord.processLiveness to the memberwise initializer, fixing a main compile break.

Written for commit dfe4185. Summary will update on new commits.

Review in cubic

Summary by CodeRabbit

  • New Features
    • Added live synchronization for device lists, relay updates, and grantless device admission.
    • Added persistent device-list state with improved sign-out cleanup.
    • Added IRX-backed Networking settings, live status updates, connection checks, and relay management.
    • Pairing QR codes now preserve Mac device identity while retaining legacy compatibility.
  • Bug Fixes
    • Pairing and token requests recover promptly after timed-out authentication.
    • Improved connectivity diagnostics, retry handling, keepalive resilience, and relay-hint updates.
    • Added a “Not verified” warning for older Mac connections.

azooz2003-bit and others added 30 commits August 26, 2026 22:14
During the 60-min irx soak (issue #10924) the iOS client failed its first
mint attempt on 15/15 autopilot cycles with journal a_error "connectivity".
The sim's unified log shows the mechanism: the mint POST reused a pooled
keep-alive connection idle since the previous ~3-4 min cycle
(reused_after_ms=171601/173084), the broker edge had closed it, the first
read returned POSIX 54 (ECONNRESET), and URLSession surfaced
NSURLErrorNetworkConnectionLost (-1005) without a transparent retry
because the POST body was already written (Apple QA1941).

The new tests drive IrxBrokerService.mintRelayCredentials through the real
URLSession stack against an in-process keep-alive HTTP server that resets
the pooled connection exactly when the second mint arrives on it. They
assert the mint survives via one immediate retry and that a final failure
carries the URL error classification; both fail on current code, which
neither retries nor attributes the error.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rors

Root cause (issue #10924): the credential autopilot mints every ~3-4
minutes, longer than the broker edge keeps an idle pooled connection
alive. The first POST of each cycle reused the dead socket
(reused_after_ms=171601/173084 in the sim's unified log), the first read
returned POSIX 54 (ECONNRESET), and URLSession surfaced
NSURLErrorNetworkConnectionLost (-1005) without a transparent retry
because the POST body was already written (Apple QA1941, CFNetwork
"idempotent(N) ... can retry(N)"). The client collapsed that to bare
`.connectivity`, so the journal could not distinguish a dead pooled
connection from real network loss, and the autopilot waited ~64s to
re-mint: 15/15 first-attempt failures over a 60-minute soak.

Fix, two pieces:

- CmxIrohTrustBrokerClientError.connectivity now carries an optional
  CmxIrohBrokerConnectivityCause with the underlying NSURLError code, and
  the error renders as e.g. `connectivity(networkConnectionLost(-1005))`,
  which the autopilot's mint-failed journaling picks up unchanged. Policy
  call sites that compared equality now use `isConnectivity`, so retry and
  cached-state classification is unaffected by the attribution.

- IrxBrokerService.mintRelayCredentials retries the idempotent bootstrap
  mint exactly once, immediately, and only for the connection-reuse class
  (.networkConnectionLost): the failed attempt already purged the dead
  pooled connection, so the retry runs on a fresh connection. The retry is
  journaled as `relay-mint-retried` with the attributed error, and the
  autopilot's half-remaining-validity sleep loop stays the outer safety
  net, unchanged.

No sleeps, no retry loops, no behavior change for any other failure class.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
schemas/control-plane/ is the single hand-editable source; quicktype
(pinned 26.0.0) generates both platforms' types via
scripts/gen-control-plane-types.sh; scripts/check-control-plane-types.sh
fails CI on drift. Fixtures are the shared golden messages both sides'
tests round-trip.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…A is bearer-authorized)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ace (both platforms)

IrxControlPlaneClient rides one WebSocket to the per-account control-plane
DO: revisioned facts in (relay passes, home-relay hints, directory),
hint announcements out (Mac). Pushed passes pass the broker's mint rules
(fleet allowlist, identity binding, monotonic freshness), rotate
make-before-break, and reset the autopilot timer so push and HTTPS
fallback never double-mint. A pushed hint that contradicts an in-flight
dial cancels it and redials immediately (IrxPeerEngine.relayHintChanged);
admitted sessions and parked denials are never touched. iOS wires it at
provisioning with the presence-service base URL; the Mac publishes its
home relay on every rotation. 20/20 package tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
# Conflicts:
#	Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxBrokerService.swift
…sence worker

Route GET /v1/control/socket: bearer-verified in the worker (same
pattern as connectivity/subscribe), then forwarded to one
AccountControlPlane DO per verified Stack user with the verified
account id and a stream deadline capped at token expiry. New DO
binding + additive v2 migration in both wrangler configs; the Vercel
origin is the optional var CMUX_WEB_BASE_URL (default
https://cmux.com). README documents the channel; tsconfig.test.json
includes the new pure modules for the bun test lane.

The DO class, protocol core, proof helper, and their tests
(controlPlane.ts, controlPlaneDo.ts, controlPlaneProof.ts,
test/controlPlane*.test.ts) were swept into merge commit 50ce4bb
by the concurrent session in this worktree; this commit adds the
remaining wiring. bun test: 228 pass (37 control-plane), typecheck
and wrangler --dry-run green for both configs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Runs scripts/check-control-plane-types.sh (quicktype regen + diff of
the committed Swift and TS outputs against schemas/control-plane/)
after the existing bun setup, so a hand edit of a generated file or a
schema change without regen fails the cheap CI layer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ker mint requires endpoint proof for non-legacy namespaces)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ant legacy stack

HostSettingsActions.irohSettingsController() always returned
MobileHostIrohRuntime.shared, which is deliberately dormant while irx owns
the transport slot, so the macOS Settings Networking section permanently
showed Inactive with stale policy info.

Route the controller by MobileHostIrxRuntime.isEnabled and give the irx
runtime its own CmxIrohSettingsControlling backend:

- runtimeStatus: relayed(home relay) when the endpoint is online (relay is
  the only attributable path in irx v1), starting during activation or an
  endpoint rebind, degraded while the activation retry ladder owns recovery,
  inactive when signed out.
- managedRelays: the trust snapshot's relay fleet with the endpoint's actual
  home relay marked selected; policySource is server only after a live
  discovery this run, cached when serving from disk; policyExpiresAt is the
  relay credentials' signed expiry.
- Unsupported mutations (relay preference narrowing, custom relays, path
  preference changes beyond the force-relay flag) throw a localized
  MobileHostIrxSettingsUnsupportedError instead of silently no-opping;
  idempotent sets that match the active state succeed.
- Snapshot stream publishes on phase changes, credential rotation, and
  pushed relay passes, plus a 30s re-yield loop that runs only while
  subscribers exist.
- Refresh forces a live broker discovery; the connection check reports the
  macHost role with relay reachability from the endpoint's relay link.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…eployment

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The broker client attaches its per-request Ed25519 binding proof only
after register() arms it on that client instance. The launch fast path
ran register in the background, so a cold launch with stale relay passes
raced a proofless mint against it, got 403
binding_request_proof_required from the hardened mint policy, and wedged
provisioning in a retry loop whose every attempt re-registered (bumping
the account route revision fleet-wide). Now: fresh cached passes keep
the fully-background zero-RTT path; stale passes register serially first
so the mint always carries proof; the provisioning retry loop backs off
2s -> 30s instead of hammering.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The web API's native auth (parseNativeStackTokens) requires BOTH
Authorization: Bearer and x-stack-refresh-token; the DO's upstream
proxy calls carried only the bearer and 401'd on every discovery/mint.
Clients now send the pair on the WS upgrade, the DO stores it per
socket (JSON value under the existing key, legacy bare-bearer values
still readable) and forwards both headers upstream. Also: upstream
refusals now log method/path/status/body-head to the worker tail, which
is how this was diagnosed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Setup no longer starts until the auth session is affirmatively
published: provisioning consumes authenticatedSessionIdentities() (first
element = current state, so signed-in launches provision immediately)
instead of polling a timer that could fire before sign-in completed and
then wait out a backoff after the session was already there. Pre-auth
attempts are gone entirely; failures past the auth gate are real broker/
network failures and keep the capped backoff.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The registry POST carries this Mac's Iroh route (the only lane that can
refresh a paired phone's stored endpoint after the Mac's identity
changes), but it rode vmAPIBaseURL, which the tag rig bakes to a
tag-local localhost origin no phone ever reads. Dev iPhones read
/api/devices from shared staging, so a phone paired to an older Mac
identity kept dialing the dead endpoint forever (observed: 143/143
dials to a stale peer across 4 launches while the Mac republished
routes only to localhost). Same split pushAPIBaseURL already fixed for
the push lane; registry publication now defaults to staging in Debug
with its own env/dev-file override, and the silent best-effort catch
now logs, so an unreachable registry is diagnosable.
The event-driven gate missed sessions restored from the keychain when
their publish predated the subscription: signed-in launches sat
unprovisioned and every transport request failed instantly with
authorization errors. Provisioning now checks the definitive state at
bootstrap completion (restored sessions provision immediately) and keeps
the identity-stream path for fresh sign-ins and account transitions.
Still zero pre-auth attempts and zero polling.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The gate's silence between 'configured' and its first outcome made two
different failure mechanisms (identity stream never delivering a
restored session vs the session snapshot hanging) indistinguishable
from the device journal. Every transition now logs: bootstrap
completion, each identity-stream delivery, snapshot check/outcome.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The launcher's device auto-pair mints an identity-only v3 pairing URL
(v=3&i=<endpoint>), which decodes into a ticket with an empty
macDeviceID. The phone's RPC client passes that as expectedPeerDeviceID
and the irx deferred transport factory correctly refuses to dial
without a named peer intent, so EVERY injected physical-device
auto-pair fails in ~30ms with missingPeerIntent and the app falls back
to its stored (possibly dead-endpoint) route. Observed live on the
phone: pairing.qr_decode.success route_kinds=iroh followed by
pairing.attempt.failed error=missingPeerIntent, then hours of dials to
a stale endpoint; likely the cause of the fleet-wide 'signed launch
failed' install-queue entries since irx became primary.

Fix: the v3 grammar gains an optional 'd=<macDeviceID>' item (encoder
writes it when present; decoder stays strict about cardinality and
unknown keys, and legacy 2-item URLs still decode with an empty id),
and the physical-device mint's decode round-trip now also proves the
device id survives. Costs one QR version (41->45 modules), still far
under the 57-module compact payload. The device id identifies but never
authorizes; admission remains the only authority.
…e perms (#11047)

Every cmux build on a Mac shared ~/Library/Application Support/cmux-irx
(Mac apps are unsandboxed), so a dev build's STAGING trust keys overwrote a
production NIGHTLY's grant-verification material: admission then denied
valid production grants as invalid-grant, making that Mac invisible to
phones (08-27 incident, second bug). The shared folder also held the
device's plaintext identity key world-readable.

State now lives at cmux-transport/<bundle-id>/<broker-host>/ (path-traversal
-safe sanitization), the legacy shared folder is deleted on sight, and
IrxDiskCache writes 0600 files in 0700 directories. Both platforms; iOS is
sandbox-isolated already but gains staging/prod separation within one app.
Keychain migration for the identity key is the follow-up.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…DiskCacheTrustReader

#11047 changed the macOS call sites to read(stateDirectory:) but left the
reader's signature as read() and gave its body references to brokerBaseURL
and stateDirectory, which only exist on MobileHostIrxRuntime. The macOS app
target has not compiled since that merge (nightly red since aa7c922).

The reader now takes the per-bundle, per-broker state directory that
activation already computes, stores, and passes at both call sites, keeping
the namespacing intent of #11047. Legacy shared-directory cleanup still
happens in activate().

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A token fetch that races the launch sign-in times out once and poisons
RPCStackTokenGate for 30s: every request inside the window throws
requestTimedOut in milliseconds. The injected dev attach fires right
after sign-in completes, lands inside that window, fails fast, and
never persists the ticket's routes, so the next cold launch dials the
stale stored endpoint again (phone log: pairing.attempt.started ->
pairing.attempt.failed error=timedOut 35ms later, while the same
endpoint admits fine 50s on when the window has lapsed).

A pairing attempt begins with fresh interactive auth, which invalidates
the suppression's premise, so it now clears the window on both token
gates before connecting. The stuck provider task is abandoned (already
cancelled; its completion watcher reaps it) and untimed-out in-flight
acquisitions are untouched.
# Conflicts:
#	Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxBrokerService.swift
#	Sources/Mobile/MobileHostIrxRuntime.swift
#	ios/cmuxPackage/Sources/cmuxFeature/MobileIrxRuntimeComposition.swift
One launch-time Stack call riding a dead pooled QUIC connection hangs
past every deadline (observed: 191s on api.stack-auth.com oauth/token),
times its phase out, and arms the 30s damper. The injected dev attach
fires seconds after the forced sign-in, lands inside that window, and
fast-fails with AuthError.timedOut in ~35ms, so the fresh pairing route
never persists and the next cold launch dials the stale stored endpoint
(the remaining 'phone never connects' leg after the ticket grammar fix;
this also poisoned the same launch's RPC token gate, fixed previously).

The damper's job is suppressing AUTOMATIC retry hammering; an explicit
user action is entitled to reclaim the phase. supersedeTimedOutAuthPhases
releases parked timed-out slots in both mechanisms (exchange registry +
token-touching states; writes stay safe via the sign-in chokepoint and
generational finishTokenTouchingPhase), and both pairing entrypoints
(open-URL and injected attach) call it before connecting. Live untimed
operations keep their exclusivity.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…s, background-return <2s)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 1 file (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="web/services/vms/images/devbox/cmux-bashrc">

<violation number="1" location="web/services/vms/images/devbox/cmux-bashrc:28">
P2: On E2B, Daytona, and Freestyle devbox images, this guard is never entered because neither build path creates `/etc/cmux/blesh-cache-seed`. Add the cache-generation, validation, and `cache.d` copy steps to the devbox Dockerfile and mirror them in the Freestyle replay before relying on this per-home seeding code.

(Based on your team's feedback about mirroring ble.sh cache generation in Freestyle replay.)</violation>
</file>

Tip: Review your code locally with the cubic CLI to iterate faster.

Re-trigger cubic

# lib/init-term.sh — ble.sh regenerates and prints unless the cache file is
# NEWER than that file), so copy each seed entry that is missing or stale.
# Plain cp (not -a) stamps now, which always beats init-term.sh.
if [ -d /etc/cmux/blesh-cache-seed/blesh ]; then

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: On E2B, Daytona, and Freestyle devbox images, this guard is never entered because neither build path creates /etc/cmux/blesh-cache-seed. Add the cache-generation, validation, and cache.d copy steps to the devbox Dockerfile and mirror them in the Freestyle replay before relying on this per-home seeding code.

(Based on your team's feedback about mirroring ble.sh cache generation in Freestyle replay.)

View Feedback

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At web/services/vms/images/devbox/cmux-bashrc, line 28:

<comment>On E2B, Daytona, and Freestyle devbox images, this guard is never entered because neither build path creates `/etc/cmux/blesh-cache-seed`. Add the cache-generation, validation, and `cache.d` copy steps to the devbox Dockerfile and mirror them in the Freestyle replay before relying on this per-home seeding code.

(Based on your team's feedback about mirroring ble.sh cache generation in Freestyle replay.) </comment>

<file context>
@@ -16,6 +16,29 @@ if [ ! -f "$HOME/.bash_history" ] && [ -f /etc/cmux/seed-history ]; then
+# lib/init-term.sh — ble.sh regenerates and prints unless the cache file is
+# NEWER than that file), so copy each seed entry that is missing or stale.
+# Plain cp (not -a) stamps now, which always beats init-term.sh.
+if [ -d /etc/cmux/blesh-cache-seed/blesh ]; then
+  __cmux_cache_base="${XDG_CACHE_HOME:-$HOME/.cache}"
+  for __cmux_seed in /etc/cmux/blesh-cache-seed/blesh/*/term.*; do
</file context>

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 2 files (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift">

<violation number="1" location="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift:137">
P2: When `URLSessionWebSocketTask.receive()` is the zombie, the timeout does not cancel the URLSession task until after the task group has waited for that receive to finish. Cancel the WebSocket from the timeout child before throwing so the timeout can actually unblock the reconnect loop.</violation>
</file>

Tip: Review your code locally with the cubic CLI to iterate faster.

Re-trigger cubic

Comment on lines +137 to +140
group.addTask {
try await Task.sleep(for: Self.receiveTimeout)
throw ReceiveTimeout()
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: When URLSessionWebSocketTask.receive() is the zombie, the timeout does not cancel the URLSession task until after the task group has waited for that receive to finish. Cancel the WebSocket from the timeout child before throwing so the timeout can actually unblock the reconnect loop.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift, line 137:

<comment>When `URLSessionWebSocketTask.receive()` is the zombie, the timeout does not cancel the URLSession task until after the task group has waited for that receive to finish. Cancel the WebSocket from the timeout child before throwing so the timeout can actually unblock the reconnect loop.</comment>

<file context>
@@ -119,6 +119,35 @@ public actor IrxControlPlaneClient {
+                of: URLSessionWebSocketTask.Message.self
+            ) { group in
+                group.addTask { try await task.receive() }
+                group.addTask {
+                    try await Task.sleep(for: Self.receiveTimeout)
+                    throw ReceiveTimeout()
</file context>
Suggested change
group.addTask {
try await Task.sleep(for: Self.receiveTimeout)
throw ReceiveTimeout()
}
group.addTask {
try await Task.sleep(for: Self.receiveTimeout)
task.cancel(with: .goingAway, reason: nil)
throw ReceiveTimeout()
}

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 5 files (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxCtlListAuthOverlays.swift">

<violation number="1" location="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxCtlListAuthOverlays.swift:6">
P2: These generated value types already satisfy synthesized `Sendable`, but `@unchecked` disables compiler verification for every future schema change. Replace the unchecked conformances with ordinary `Sendable` conformances so a newly generated non-Sendable field fails at compile time instead of crossing the actor boundary unsafely.</violation>
</file>

Tip: Review your code locally with the cubic CLI to iterate faster.

Re-trigger cubic

Comment on lines +6 to +12
extension Binding: @unchecked Sendable {}
extension CTLDirectory: @unchecked Sendable {}
extension CTLDirectoryPayload: @unchecked Sendable {}
extension GrantVerificationKey: @unchecked Sendable {}
extension PurpleMinimumSupportedVersion: @unchecked Sendable {}
extension ReleaseTrack: @unchecked Sendable {}
extension Status: @unchecked Sendable {}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: These generated value types already satisfy synthesized Sendable, but @unchecked disables compiler verification for every future schema change. Replace the unchecked conformances with ordinary Sendable conformances so a newly generated non-Sendable field fails at compile time instead of crossing the actor boundary unsafely.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxCtlListAuthOverlays.swift, line 6:

<comment>These generated value types already satisfy synthesized `Sendable`, but `@unchecked` disables compiler verification for every future schema change. Replace the unchecked conformances with ordinary `Sendable` conformances so a newly generated non-Sendable field fails at compile time instead of crossing the actor boundary unsafely.</comment>

<file context>
@@ -1,9 +1,15 @@
+// Generated wire values contain only immutable value types; make that fact
+// explicit so the actor's handler closures can receive decoded directory
+// payloads without changing the schema-generated source.
+extension Binding: @unchecked Sendable {}
+extension CTLDirectory: @unchecked Sendable {}
+extension CTLDirectoryPayload: @unchecked Sendable {}
</file context>
Suggested change
extension Binding: @unchecked Sendable {}
extension CTLDirectory: @unchecked Sendable {}
extension CTLDirectoryPayload: @unchecked Sendable {}
extension GrantVerificationKey: @unchecked Sendable {}
extension PurpleMinimumSupportedVersion: @unchecked Sendable {}
extension ReleaseTrack: @unchecked Sendable {}
extension Status: @unchecked Sendable {}
extension Binding: Sendable {}
extension CTLDirectory: Sendable {}
extension CTLDirectoryPayload: Sendable {}
extension GrantVerificationKey: Sendable {}
extension PurpleMinimumSupportedVersion: Sendable {}
extension ReleaseTrack: Sendable {}
extension Status: Sendable {}

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 6 files (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift">

<violation number="1" location="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift:364">
P2: When a relay-pass frame arrives while a directory revision is pending, this ACK marks the directory revision applied even though `onRelayPasses` does not apply the device list. The DO can therefore clear its directory retry ladder and leave admission state stale; acknowledge only after applying a directory/hint fact, or track relay-pass acknowledgements separately.</violation>
</file>

Tip: Review your code locally with the cubic CLI to iterate faster.

Re-trigger cubic

Comment thread Sources/Mobile/MobileHostIrxRuntime.swift
["rev": String(fact.rev), "count": String(credentials.count)]
)
await handlers.onRelayPasses(credentials)
// Relay credentials are a revisioned control-plane fact. The

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: When a relay-pass frame arrives while a directory revision is pending, this ACK marks the directory revision applied even though onRelayPasses does not apply the device list. The DO can therefore clear its directory retry ladder and leave admission state stale; acknowledge only after applying a directory/hint fact, or track relay-pass acknowledgements separately.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift, line 364:

<comment>When a relay-pass frame arrives while a directory revision is pending, this ACK marks the directory revision applied even though `onRelayPasses` does not apply the device list. The DO can therefore clear its directory retry ladder and leave admission state stale; acknowledge only after applying a directory/hint fact, or track relay-pass acknowledgements separately.</comment>

<file context>
@@ -361,6 +361,10 @@ public actor IrxControlPlaneClient {
                     ["rev": String(fact.rev), "count": String(credentials.count)]
                 )
                 await handlers.onRelayPasses(credentials)
+                // Relay credentials are a revisioned control-plane fact. The
+                // handler has completed, so acknowledge it now to stop the
+                // server retry ladder and advance convergence state.
</file context>

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 3 files (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="Sources/Mobile/MobileHostIrxRuntime.swift">

<violation number="1" location="Sources/Mobile/MobileHostIrxRuntime.swift:658">
P2: When the list is stale, absent, or the endpoint was delisted between the hello and registry check, this callback returns `false` and the registry reports `.revoked` for every case. Preserve the judge's denial code through the registry check so recoverable lease/list failures are not presented as permanent revocation.</violation>
</file>

Tip: Review your code locally with the cubic CLI to iterate faster.

Re-trigger cubic

deviceID: peer.deviceID,
sessionID: sessionID,
connection: irx,
stillAuthorized: { endpointIDHex in

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: When the list is stale, absent, or the endpoint was delisted between the hello and registry check, this callback returns false and the registry reports .revoked for every case. Preserve the judge's denial code through the registry check so recoverable lease/list failures are not presented as permanent revocation.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Sources/Mobile/MobileHostIrxRuntime.swift, line 658:

<comment>When the list is stale, absent, or the endpoint was delisted between the hello and registry check, this callback returns `false` and the registry reports `.revoked` for every case. Preserve the judge's denial code through the registry check so recoverable lease/list failures are not presented as permanent revocation.</comment>

<file context>
@@ -651,7 +651,20 @@ final class MobileHostIrxRuntime {
+            deviceID: peer.deviceID,
+            sessionID: sessionID,
+            connection: irx,
+            stillAuthorized: { endpointIDHex in
+                do {
+                    _ = try judge.judgment()(nil, endpointIDHex)
</file context>

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 2 files (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxJSONCache.swift">

<violation number="1" location="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxJSONCache.swift:176">
P2: Release construction now preserves legacy cache files without performing the identity-validated migration promised by this comment. Wire the migration and cleanup into the factory, or remove the legacy files only after a validated replacement, so stale account state cannot persist and be reused by an unscoped broker.</violation>
</file>

Tip: Review your code locally with the cubic CLI to iterate faster.

Re-trigger cubic

guard let scope else { return file }
// A legacy snapshot is not account/backend scoped. Never import it
// into the scoped keychain item, even when its shape happens to decode.
// Preserve the old copy until an explicit, identity-validated

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: Release construction now preserves legacy cache files without performing the identity-validated migration promised by this comment. Wire the migration and cleanup into the factory, or remove the legacy files only after a validated replacement, so stale account state cannot persist and be reused by an unscoped broker.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxJSONCache.swift, line 176:

<comment>Release construction now preserves legacy cache files without performing the identity-validated migration promised by this comment. Wire the migration and cleanup into the factory, or remove the legacy files only after a validated replacement, so stale account state cannot persist and be reused by an unscoped broker.</comment>

<file context>
@@ -173,9 +173,9 @@ enum IrxBrokerCacheFactory {
-        // Remove the old copy so a later unscoped bootstrap cannot resurrect
-        // credentials after sign-out or account switching.
-        file.clear()
+        // Preserve the old copy until an explicit, identity-validated
+        // migration can replace it; construction must never destroy the only
+        // offline binding or credential state during an upgrade.
</file context>

@cursor

cursor Bot commented Sep 1, 2026

Copy link
Copy Markdown

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 issues found across 3 files (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift">

<violation number="1" location="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift:314">
P1: When `kick()` or `stop()` runs while a frame handler is suspended, the old `route` continuation still applies that frame after the new generation starts. Recheck the generation after each awaited handler and before acknowledging, and bind acknowledgements to the frame's socket so stale frames cannot mutate the new control-plane state.</violation>
</file>

<file name="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxBrokerService.swift">

<violation number="1" location="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxBrokerService.swift:514">
P2: When a cached pushed snapshot contains duplicate relay URLs, `Dictionary(uniqueKeysWithValues:)` traps before the next pass is processed. Build the map with a uniquing closure so malformed or legacy cache data is coalesced safely.</violation>
</file>

Tip: Review your code locally with the cubic CLI to iterate faster.

Re-trigger cubic

case .data(let raw): data = raw
@unknown default: continue
}
await route(data, generation: generation)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: When kick() or stop() runs while a frame handler is suspended, the old route continuation still applies that frame after the new generation starts. Recheck the generation after each awaited handler and before acknowledging, and bind acknowledgements to the frame's socket so stale frames cannot mutate the new control-plane state.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift, line 314:

<comment>When `kick()` or `stop()` runs while a frame handler is suspended, the old `route` continuation still applies that frame after the new generation starts. Recheck the generation after each awaited handler and before acknowledging, and bind acknowledgements to the frame's socket so stale frames cannot mutate the new control-plane state.</comment>

<file context>
@@ -290,19 +302,21 @@ public actor IrxControlPlaneClient {
             @unknown default: continue
             }
-            await route(data)
+            await route(data, generation: generation)
         }
     }
</file context>

Comment on lines +514 to +517
var mergedByRelay = Dictionary(
uniqueKeysWithValues: (cachedSnapshot?.credentials ?? []).map {
($0.relayURL, $0)
})

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: When a cached pushed snapshot contains duplicate relay URLs, Dictionary(uniqueKeysWithValues:) traps before the next pass is processed. Build the map with a uniquing closure so malformed or legacy cache data is coalesced safely.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxBrokerService.swift, line 514:

<comment>When a cached pushed snapshot contains duplicate relay URLs, `Dictionary(uniqueKeysWithValues:)` traps before the next pass is processed. Build the map with a uniquing closure so malformed or legacy cache data is coalesced safely.</comment>

<file context>
@@ -506,28 +508,44 @@ public actor IrxBrokerService {
+        let cachedSnapshot = credentialCache.load().flatMap { snapshot in
+            snapshot.endpointIDHex == identity.endpointIDHex ? snapshot : nil
+        }
+        var mergedByRelay = Dictionary(
+            uniqueKeysWithValues: (cachedSnapshot?.credentials ?? []).map {
+                ($0.relayURL, $0)
</file context>
Suggested change
var mergedByRelay = Dictionary(
uniqueKeysWithValues: (cachedSnapshot?.credentials ?? []).map {
($0.relayURL, $0)
})
var mergedByRelay = Dictionary(
(cachedSnapshot?.credentials ?? []).map {
($0.relayURL, $0)
},
uniquingKeysWith: { _, latest in latest })

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

3 issues found across 5 files (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift">

<violation number="1" location="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift:431">
P1: When a directory callback fails during an initial snapshot, this line suppresses the ack but `snapshot_complete` still persists that revision as `haveRev`. The next reconnect takes the DO's resumed fast path and skips the failed directory, leaving list-auth state stale; persist the cursor only through the highest contiguous applied revision, or otherwise force an unacknowledged snapshot to replay.</violation>
</file>

<file name="Sources/Mobile/MobileHostIrxRuntime.swift">

<violation number="1" location="Sources/Mobile/MobileHostIrxRuntime.swift:470">
P1: When the secure-store write fails, this guard drops the new directory before enforcing it, so a revoked or delisted peer remains admitted under the old lease. Apply and enforce the snapshot in memory (or clear admission fail-closed) even when persistence fails, while returning `false` so the server retries instead of acknowledging it.</violation>
</file>

<file name="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxBrokerService.swift">

<violation number="1" location="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxBrokerService.swift:521">
P3: When the new guard rejects pushed credentials because they are already expired (or have refreshAfter >= expiresAt), the journal records reason "stale" in the `!accepted.isEmpty` fallback, the same label already used for monotonic-freshness drops. That conflates the two rejection modes and will mislead anyone debugging the replay-rejection path. Record a distinct reason (e.g. "expired") for the new guard's rejections.</violation>
</file>

Tip: Review your code locally with the cubic CLI to iterate faster.

Re-trigger cubic

}
// Directory delivery is itself a revisioned fact. Ack only
// after every consumer reports durable application.
if applied { await acknowledge(rev: listFact.rev) }

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: When a directory callback fails during an initial snapshot, this line suppresses the ack but snapshot_complete still persists that revision as haveRev. The next reconnect takes the DO's resumed fast path and skips the failed directory, leaving list-auth state stale; persist the cursor only through the highest contiguous applied revision, or otherwise force an unacknowledged snapshot to replay.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift, line 431:

<comment>When a directory callback fails during an initial snapshot, this line suppresses the ack but `snapshot_complete` still persists that revision as `haveRev`. The next reconnect takes the DO's resumed fast path and skips the failed directory, leaving list-auth state stale; persist the cursor only through the highest contiguous applied revision, or otherwise force an unacknowledged snapshot to replay.</comment>

<file context>
@@ -410,26 +411,24 @@ public actor IrxControlPlaneClient {
-                await acknowledge(rev: listFact.rev)
+                // Directory delivery is itself a revisioned fact. Ack only
+                // after every consumer reports durable application.
+                if applied { await acknowledge(rev: listFact.rev) }
             case "current":
                 // Explicit freshness re-stamp for the device-list lease.
</file context>

receivedAtWall: Date(),
receivedAtMonotonic: .now
)
guard await deviceListStore.persist(snapshot) else { return false }

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: When the secure-store write fails, this guard drops the new directory before enforcing it, so a revoked or delisted peer remains admitted under the old lease. Apply and enforce the snapshot in memory (or clear admission fail-closed) even when persistence fails, while returning false so the server retries instead of acknowledging it.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Sources/Mobile/MobileHostIrxRuntime.swift, line 470:

<comment>When the secure-store write fails, this guard drops the new directory before enforcing it, so a revoked or delisted peer remains admitted under the old lease. Apply and enforce the snapshot in memory (or clear admission fail-closed) even when persistence fails, while returning `false` so the server retries instead of acknowledging it.</comment>

<file context>
@@ -452,23 +453,22 @@ final class MobileHostIrxRuntime {
             receivedAtWall: Date(),
             receivedAtMonotonic: .now
         )
+        guard await deviceListStore.persist(snapshot) else { return false }
         deviceListBox.replace(snapshot)
-        await deviceListStore.persist(snapshot)
</file context>

var accepted: [IrxRelayCredential] = []
let now = Date()
for credential in pushed {
guard credential.expiresAt > now,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: When the new guard rejects pushed credentials because they are already expired (or have refreshAfter >= expiresAt), the journal records reason "stale" in the !accepted.isEmpty fallback, the same label already used for monotonic-freshness drops. That conflates the two rejection modes and will mislead anyone debugging the replay-rejection path. Record a distinct reason (e.g. "expired") for the new guard's rejections.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxBrokerService.swift, line 521:

<comment>When the new guard rejects pushed credentials because they are already expired (or have refreshAfter >= expiresAt), the journal records reason "stale" in the `!accepted.isEmpty` fallback, the same label already used for monotonic-freshness drops. That conflates the two rejection modes and will mislead anyone debugging the replay-rejection path. Record a distinct reason (e.g. "expired") for the new guard's rejections.</comment>

<file context>
@@ -516,7 +516,11 @@ public actor IrxBrokerService {
         var accepted: [IrxRelayCredential] = []
+        let now = Date()
         for credential in pushed {
+            guard credential.expiresAt > now,
+                credential.refreshAfter < credential.expiresAt
+            else { continue }
</file context>

@azooz2003-bit
azooz2003-bit merged commit 9bbf2fe into main Sep 1, 2026
17 of 21 checks passed
@azooz2003-bit
azooz2003-bit deleted the feat-listauth branch September 1, 2026 07:38
@github-actions github-actions Bot locked as resolved and limited conversation to collaborators Sep 1, 2026

This branch was successfully deployed

2 active deployments
Preview – cmux166 — dfe4185a Deployed Sep 1, 2026 by vercel[bot]
Preview – cmux41 — dfe4185a Deployed Sep 1, 2026 by vercel[bot]
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant