Repository navigation
Device-list authorization + DO-first control plane (listauth) - #11220
Conversation
During the 60-min irx soak (issue #10924) the iOS client failed its first mint attempt on 15/15 autopilot cycles with journal a_error "connectivity". The sim's unified log shows the mechanism: the mint POST reused a pooled keep-alive connection idle since the previous ~3-4 min cycle (reused_after_ms=171601/173084), the broker edge had closed it, the first read returned POSIX 54 (ECONNRESET), and URLSession surfaced NSURLErrorNetworkConnectionLost (-1005) without a transparent retry because the POST body was already written (Apple QA1941). The new tests drive IrxBrokerService.mintRelayCredentials through the real URLSession stack against an in-process keep-alive HTTP server that resets the pooled connection exactly when the second mint arrives on it. They assert the mint survives via one immediate retry and that a final failure carries the URL error classification; both fail on current code, which neither retries nor attributes the error. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rors Root cause (issue #10924): the credential autopilot mints every ~3-4 minutes, longer than the broker edge keeps an idle pooled connection alive. The first POST of each cycle reused the dead socket (reused_after_ms=171601/173084 in the sim's unified log), the first read returned POSIX 54 (ECONNRESET), and URLSession surfaced NSURLErrorNetworkConnectionLost (-1005) without a transparent retry because the POST body was already written (Apple QA1941, CFNetwork "idempotent(N) ... can retry(N)"). The client collapsed that to bare `.connectivity`, so the journal could not distinguish a dead pooled connection from real network loss, and the autopilot waited ~64s to re-mint: 15/15 first-attempt failures over a 60-minute soak. Fix, two pieces: - CmxIrohTrustBrokerClientError.connectivity now carries an optional CmxIrohBrokerConnectivityCause with the underlying NSURLError code, and the error renders as e.g. `connectivity(networkConnectionLost(-1005))`, which the autopilot's mint-failed journaling picks up unchanged. Policy call sites that compared equality now use `isConnectivity`, so retry and cached-state classification is unaffected by the attribution. - IrxBrokerService.mintRelayCredentials retries the idempotent bootstrap mint exactly once, immediately, and only for the connection-reuse class (.networkConnectionLost): the failed attempt already purged the dead pooled connection, so the retry runs on a fresh connection. The retry is journaled as `relay-mint-retried` with the attributed error, and the autopilot's half-remaining-validity sleep loop stays the outer safety net, unchanged. No sleeps, no retry loops, no behavior change for any other failure class. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
schemas/control-plane/ is the single hand-editable source; quicktype (pinned 26.0.0) generates both platforms' types via scripts/gen-control-plane-types.sh; scripts/check-control-plane-types.sh fails CI on drift. Fixtures are the shared golden messages both sides' tests round-trip. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…A is bearer-authorized) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ace (both platforms) IrxControlPlaneClient rides one WebSocket to the per-account control-plane DO: revisioned facts in (relay passes, home-relay hints, directory), hint announcements out (Mac). Pushed passes pass the broker's mint rules (fleet allowlist, identity binding, monotonic freshness), rotate make-before-break, and reset the autopilot timer so push and HTTPS fallback never double-mint. A pushed hint that contradicts an in-flight dial cancels it and redials immediately (IrxPeerEngine.relayHintChanged); admitted sessions and parked denials are never touched. iOS wires it at provisioning with the presence-service base URL; the Mac publishes its home relay on every rotation. 20/20 package tests green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
# Conflicts: # Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxBrokerService.swift
…sence worker Route GET /v1/control/socket: bearer-verified in the worker (same pattern as connectivity/subscribe), then forwarded to one AccountControlPlane DO per verified Stack user with the verified account id and a stream deadline capped at token expiry. New DO binding + additive v2 migration in both wrangler configs; the Vercel origin is the optional var CMUX_WEB_BASE_URL (default https://cmux.com). README documents the channel; tsconfig.test.json includes the new pure modules for the bun test lane. The DO class, protocol core, proof helper, and their tests (controlPlane.ts, controlPlaneDo.ts, controlPlaneProof.ts, test/controlPlane*.test.ts) were swept into merge commit 50ce4bb by the concurrent session in this worktree; this commit adds the remaining wiring. bun test: 228 pass (37 control-plane), typecheck and wrangler --dry-run green for both configs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Runs scripts/check-control-plane-types.sh (quicktype regen + diff of the committed Swift and TS outputs against schemas/control-plane/) after the existing bun setup, so a hand edit of a generated file or a schema change without regen fails the cheap CI layer. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ker mint requires endpoint proof for non-legacy namespaces) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ant legacy stack HostSettingsActions.irohSettingsController() always returned MobileHostIrohRuntime.shared, which is deliberately dormant while irx owns the transport slot, so the macOS Settings Networking section permanently showed Inactive with stale policy info. Route the controller by MobileHostIrxRuntime.isEnabled and give the irx runtime its own CmxIrohSettingsControlling backend: - runtimeStatus: relayed(home relay) when the endpoint is online (relay is the only attributable path in irx v1), starting during activation or an endpoint rebind, degraded while the activation retry ladder owns recovery, inactive when signed out. - managedRelays: the trust snapshot's relay fleet with the endpoint's actual home relay marked selected; policySource is server only after a live discovery this run, cached when serving from disk; policyExpiresAt is the relay credentials' signed expiry. - Unsupported mutations (relay preference narrowing, custom relays, path preference changes beyond the force-relay flag) throw a localized MobileHostIrxSettingsUnsupportedError instead of silently no-opping; idempotent sets that match the active state succeed. - Snapshot stream publishes on phase changes, credential rotation, and pushed relay passes, plus a 30s re-yield loop that runs only while subscribers exist. - Refresh forces a live broker discovery; the connection check reports the macHost role with relay reachability from the endpoint's relay link. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…eployment Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The broker client attaches its per-request Ed25519 binding proof only after register() arms it on that client instance. The launch fast path ran register in the background, so a cold launch with stale relay passes raced a proofless mint against it, got 403 binding_request_proof_required from the hardened mint policy, and wedged provisioning in a retry loop whose every attempt re-registered (bumping the account route revision fleet-wide). Now: fresh cached passes keep the fully-background zero-RTT path; stale passes register serially first so the mint always carries proof; the provisioning retry loop backs off 2s -> 30s instead of hammering. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The web API's native auth (parseNativeStackTokens) requires BOTH Authorization: Bearer and x-stack-refresh-token; the DO's upstream proxy calls carried only the bearer and 401'd on every discovery/mint. Clients now send the pair on the WS upgrade, the DO stores it per socket (JSON value under the existing key, legacy bare-bearer values still readable) and forwards both headers upstream. Also: upstream refusals now log method/path/status/body-head to the worker tail, which is how this was diagnosed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Setup no longer starts until the auth session is affirmatively published: provisioning consumes authenticatedSessionIdentities() (first element = current state, so signed-in launches provision immediately) instead of polling a timer that could fire before sign-in completed and then wait out a backoff after the session was already there. Pre-auth attempts are gone entirely; failures past the auth gate are real broker/ network failures and keep the capped backoff. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The registry POST carries this Mac's Iroh route (the only lane that can refresh a paired phone's stored endpoint after the Mac's identity changes), but it rode vmAPIBaseURL, which the tag rig bakes to a tag-local localhost origin no phone ever reads. Dev iPhones read /api/devices from shared staging, so a phone paired to an older Mac identity kept dialing the dead endpoint forever (observed: 143/143 dials to a stale peer across 4 launches while the Mac republished routes only to localhost). Same split pushAPIBaseURL already fixed for the push lane; registry publication now defaults to staging in Debug with its own env/dev-file override, and the silent best-effort catch now logs, so an unreachable registry is diagnosable.
The event-driven gate missed sessions restored from the keychain when their publish predated the subscription: signed-in launches sat unprovisioned and every transport request failed instantly with authorization errors. Provisioning now checks the definitive state at bootstrap completion (restored sessions provision immediately) and keeps the identity-stream path for fresh sign-ins and account transitions. Still zero pre-auth attempts and zero polling. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The gate's silence between 'configured' and its first outcome made two different failure mechanisms (identity stream never delivering a restored session vs the session snapshot hanging) indistinguishable from the device journal. Every transition now logs: bootstrap completion, each identity-stream delivery, snapshot check/outcome. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The launcher's device auto-pair mints an identity-only v3 pairing URL (v=3&i=<endpoint>), which decodes into a ticket with an empty macDeviceID. The phone's RPC client passes that as expectedPeerDeviceID and the irx deferred transport factory correctly refuses to dial without a named peer intent, so EVERY injected physical-device auto-pair fails in ~30ms with missingPeerIntent and the app falls back to its stored (possibly dead-endpoint) route. Observed live on the phone: pairing.qr_decode.success route_kinds=iroh followed by pairing.attempt.failed error=missingPeerIntent, then hours of dials to a stale endpoint; likely the cause of the fleet-wide 'signed launch failed' install-queue entries since irx became primary. Fix: the v3 grammar gains an optional 'd=<macDeviceID>' item (encoder writes it when present; decoder stays strict about cardinality and unknown keys, and legacy 2-item URLs still decode with an empty id), and the physical-device mint's decode round-trip now also proves the device id survives. Costs one QR version (41->45 modules), still far under the 57-module compact payload. The device id identifies but never authorizes; admission remains the only authority.
…e perms (#11047) Every cmux build on a Mac shared ~/Library/Application Support/cmux-irx (Mac apps are unsandboxed), so a dev build's STAGING trust keys overwrote a production NIGHTLY's grant-verification material: admission then denied valid production grants as invalid-grant, making that Mac invisible to phones (08-27 incident, second bug). The shared folder also held the device's plaintext identity key world-readable. State now lives at cmux-transport/<bundle-id>/<broker-host>/ (path-traversal -safe sanitization), the legacy shared folder is deleted on sight, and IrxDiskCache writes 0600 files in 0700 directories. Both platforms; iOS is sandbox-isolated already but gains staging/prod separation within one app. Keychain migration for the identity key is the follow-up. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…DiskCacheTrustReader #11047 changed the macOS call sites to read(stateDirectory:) but left the reader's signature as read() and gave its body references to brokerBaseURL and stateDirectory, which only exist on MobileHostIrxRuntime. The macOS app target has not compiled since that merge (nightly red since aa7c922). The reader now takes the per-bundle, per-broker state directory that activation already computes, stores, and passes at both call sites, keeping the namespacing intent of #11047. Legacy shared-directory cleanup still happens in activate(). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A token fetch that races the launch sign-in times out once and poisons RPCStackTokenGate for 30s: every request inside the window throws requestTimedOut in milliseconds. The injected dev attach fires right after sign-in completes, lands inside that window, fails fast, and never persists the ticket's routes, so the next cold launch dials the stale stored endpoint again (phone log: pairing.attempt.started -> pairing.attempt.failed error=timedOut 35ms later, while the same endpoint admits fine 50s on when the window has lapsed). A pairing attempt begins with fresh interactive auth, which invalidates the suppression's premise, so it now clears the window on both token gates before connecting. The stuck provider task is abandoned (already cancelled; its completion watcher reaps it) and untimed-out in-flight acquisitions are untouched.
# Conflicts: # Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxBrokerService.swift # Sources/Mobile/MobileHostIrxRuntime.swift # ios/cmuxPackage/Sources/cmuxFeature/MobileIrxRuntimeComposition.swift
One launch-time Stack call riding a dead pooled QUIC connection hangs past every deadline (observed: 191s on api.stack-auth.com oauth/token), times its phase out, and arms the 30s damper. The injected dev attach fires seconds after the forced sign-in, lands inside that window, and fast-fails with AuthError.timedOut in ~35ms, so the fresh pairing route never persists and the next cold launch dials the stale stored endpoint (the remaining 'phone never connects' leg after the ticket grammar fix; this also poisoned the same launch's RPC token gate, fixed previously). The damper's job is suppressing AUTOMATIC retry hammering; an explicit user action is entitled to reclaim the phase. supersedeTimedOutAuthPhases releases parked timed-out slots in both mechanisms (exchange registry + token-touching states; writes stay safe via the sign-in chokepoint and generational finishTokenTouchingPhase), and both pairing entrypoints (open-URL and injected attach) call it before connecting. Live untimed operations keep their exclusivity.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…s, background-return <2s) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
There was a problem hiding this comment.
1 issue found across 1 file (changes from recent commits).
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="web/services/vms/images/devbox/cmux-bashrc">
<violation number="1" location="web/services/vms/images/devbox/cmux-bashrc:28">
P2: On E2B, Daytona, and Freestyle devbox images, this guard is never entered because neither build path creates `/etc/cmux/blesh-cache-seed`. Add the cache-generation, validation, and `cache.d` copy steps to the devbox Dockerfile and mirror them in the Freestyle replay before relying on this per-home seeding code.
(Based on your team's feedback about mirroring ble.sh cache generation in Freestyle replay.)</violation>
</file>
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
| # lib/init-term.sh — ble.sh regenerates and prints unless the cache file is | ||
| # NEWER than that file), so copy each seed entry that is missing or stale. | ||
| # Plain cp (not -a) stamps now, which always beats init-term.sh. | ||
| if [ -d /etc/cmux/blesh-cache-seed/blesh ]; then |
There was a problem hiding this comment.
P2: On E2B, Daytona, and Freestyle devbox images, this guard is never entered because neither build path creates /etc/cmux/blesh-cache-seed. Add the cache-generation, validation, and cache.d copy steps to the devbox Dockerfile and mirror them in the Freestyle replay before relying on this per-home seeding code.
(Based on your team's feedback about mirroring ble.sh cache generation in Freestyle replay.)
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At web/services/vms/images/devbox/cmux-bashrc, line 28:
<comment>On E2B, Daytona, and Freestyle devbox images, this guard is never entered because neither build path creates `/etc/cmux/blesh-cache-seed`. Add the cache-generation, validation, and `cache.d` copy steps to the devbox Dockerfile and mirror them in the Freestyle replay before relying on this per-home seeding code.
(Based on your team's feedback about mirroring ble.sh cache generation in Freestyle replay.) </comment>
<file context>
@@ -16,6 +16,29 @@ if [ ! -f "$HOME/.bash_history" ] && [ -f /etc/cmux/seed-history ]; then
+# lib/init-term.sh — ble.sh regenerates and prints unless the cache file is
+# NEWER than that file), so copy each seed entry that is missing or stale.
+# Plain cp (not -a) stamps now, which always beats init-term.sh.
+if [ -d /etc/cmux/blesh-cache-seed/blesh ]; then
+ __cmux_cache_base="${XDG_CACHE_HOME:-$HOME/.cache}"
+ for __cmux_seed in /etc/cmux/blesh-cache-seed/blesh/*/term.*; do
</file context>
There was a problem hiding this comment.
1 issue found across 2 files (changes from recent commits).
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift">
<violation number="1" location="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift:137">
P2: When `URLSessionWebSocketTask.receive()` is the zombie, the timeout does not cancel the URLSession task until after the task group has waited for that receive to finish. Cancel the WebSocket from the timeout child before throwing so the timeout can actually unblock the reconnect loop.</violation>
</file>
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
| group.addTask { | ||
| try await Task.sleep(for: Self.receiveTimeout) | ||
| throw ReceiveTimeout() | ||
| } |
There was a problem hiding this comment.
P2: When URLSessionWebSocketTask.receive() is the zombie, the timeout does not cancel the URLSession task until after the task group has waited for that receive to finish. Cancel the WebSocket from the timeout child before throwing so the timeout can actually unblock the reconnect loop.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift, line 137:
<comment>When `URLSessionWebSocketTask.receive()` is the zombie, the timeout does not cancel the URLSession task until after the task group has waited for that receive to finish. Cancel the WebSocket from the timeout child before throwing so the timeout can actually unblock the reconnect loop.</comment>
<file context>
@@ -119,6 +119,35 @@ public actor IrxControlPlaneClient {
+ of: URLSessionWebSocketTask.Message.self
+ ) { group in
+ group.addTask { try await task.receive() }
+ group.addTask {
+ try await Task.sleep(for: Self.receiveTimeout)
+ throw ReceiveTimeout()
</file context>
| group.addTask { | |
| try await Task.sleep(for: Self.receiveTimeout) | |
| throw ReceiveTimeout() | |
| } | |
| group.addTask { | |
| try await Task.sleep(for: Self.receiveTimeout) | |
| task.cancel(with: .goingAway, reason: nil) | |
| throw ReceiveTimeout() | |
| } |
There was a problem hiding this comment.
1 issue found across 5 files (changes from recent commits).
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxCtlListAuthOverlays.swift">
<violation number="1" location="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxCtlListAuthOverlays.swift:6">
P2: These generated value types already satisfy synthesized `Sendable`, but `@unchecked` disables compiler verification for every future schema change. Replace the unchecked conformances with ordinary `Sendable` conformances so a newly generated non-Sendable field fails at compile time instead of crossing the actor boundary unsafely.</violation>
</file>
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
| extension Binding: @unchecked Sendable {} | ||
| extension CTLDirectory: @unchecked Sendable {} | ||
| extension CTLDirectoryPayload: @unchecked Sendable {} | ||
| extension GrantVerificationKey: @unchecked Sendable {} | ||
| extension PurpleMinimumSupportedVersion: @unchecked Sendable {} | ||
| extension ReleaseTrack: @unchecked Sendable {} | ||
| extension Status: @unchecked Sendable {} |
There was a problem hiding this comment.
P2: These generated value types already satisfy synthesized Sendable, but @unchecked disables compiler verification for every future schema change. Replace the unchecked conformances with ordinary Sendable conformances so a newly generated non-Sendable field fails at compile time instead of crossing the actor boundary unsafely.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxCtlListAuthOverlays.swift, line 6:
<comment>These generated value types already satisfy synthesized `Sendable`, but `@unchecked` disables compiler verification for every future schema change. Replace the unchecked conformances with ordinary `Sendable` conformances so a newly generated non-Sendable field fails at compile time instead of crossing the actor boundary unsafely.</comment>
<file context>
@@ -1,9 +1,15 @@
+// Generated wire values contain only immutable value types; make that fact
+// explicit so the actor's handler closures can receive decoded directory
+// payloads without changing the schema-generated source.
+extension Binding: @unchecked Sendable {}
+extension CTLDirectory: @unchecked Sendable {}
+extension CTLDirectoryPayload: @unchecked Sendable {}
</file context>
| extension Binding: @unchecked Sendable {} | |
| extension CTLDirectory: @unchecked Sendable {} | |
| extension CTLDirectoryPayload: @unchecked Sendable {} | |
| extension GrantVerificationKey: @unchecked Sendable {} | |
| extension PurpleMinimumSupportedVersion: @unchecked Sendable {} | |
| extension ReleaseTrack: @unchecked Sendable {} | |
| extension Status: @unchecked Sendable {} | |
| extension Binding: Sendable {} | |
| extension CTLDirectory: Sendable {} | |
| extension CTLDirectoryPayload: Sendable {} | |
| extension GrantVerificationKey: Sendable {} | |
| extension PurpleMinimumSupportedVersion: Sendable {} | |
| extension ReleaseTrack: Sendable {} | |
| extension Status: Sendable {} |
There was a problem hiding this comment.
1 issue found across 6 files (changes from recent commits).
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift">
<violation number="1" location="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift:364">
P2: When a relay-pass frame arrives while a directory revision is pending, this ACK marks the directory revision applied even though `onRelayPasses` does not apply the device list. The DO can therefore clear its directory retry ladder and leave admission state stale; acknowledge only after applying a directory/hint fact, or track relay-pass acknowledgements separately.</violation>
</file>
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
| ["rev": String(fact.rev), "count": String(credentials.count)] | ||
| ) | ||
| await handlers.onRelayPasses(credentials) | ||
| // Relay credentials are a revisioned control-plane fact. The |
There was a problem hiding this comment.
P2: When a relay-pass frame arrives while a directory revision is pending, this ACK marks the directory revision applied even though onRelayPasses does not apply the device list. The DO can therefore clear its directory retry ladder and leave admission state stale; acknowledge only after applying a directory/hint fact, or track relay-pass acknowledgements separately.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift, line 364:
<comment>When a relay-pass frame arrives while a directory revision is pending, this ACK marks the directory revision applied even though `onRelayPasses` does not apply the device list. The DO can therefore clear its directory retry ladder and leave admission state stale; acknowledge only after applying a directory/hint fact, or track relay-pass acknowledgements separately.</comment>
<file context>
@@ -361,6 +361,10 @@ public actor IrxControlPlaneClient {
["rev": String(fact.rev), "count": String(credentials.count)]
)
await handlers.onRelayPasses(credentials)
+ // Relay credentials are a revisioned control-plane fact. The
+ // handler has completed, so acknowledge it now to stop the
+ // server retry ladder and advance convergence state.
</file context>
There was a problem hiding this comment.
1 issue found across 3 files (changes from recent commits).
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="Sources/Mobile/MobileHostIrxRuntime.swift">
<violation number="1" location="Sources/Mobile/MobileHostIrxRuntime.swift:658">
P2: When the list is stale, absent, or the endpoint was delisted between the hello and registry check, this callback returns `false` and the registry reports `.revoked` for every case. Preserve the judge's denial code through the registry check so recoverable lease/list failures are not presented as permanent revocation.</violation>
</file>
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
| deviceID: peer.deviceID, | ||
| sessionID: sessionID, | ||
| connection: irx, | ||
| stillAuthorized: { endpointIDHex in |
There was a problem hiding this comment.
P2: When the list is stale, absent, or the endpoint was delisted between the hello and registry check, this callback returns false and the registry reports .revoked for every case. Preserve the judge's denial code through the registry check so recoverable lease/list failures are not presented as permanent revocation.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Sources/Mobile/MobileHostIrxRuntime.swift, line 658:
<comment>When the list is stale, absent, or the endpoint was delisted between the hello and registry check, this callback returns `false` and the registry reports `.revoked` for every case. Preserve the judge's denial code through the registry check so recoverable lease/list failures are not presented as permanent revocation.</comment>
<file context>
@@ -651,7 +651,20 @@ final class MobileHostIrxRuntime {
+ deviceID: peer.deviceID,
+ sessionID: sessionID,
+ connection: irx,
+ stillAuthorized: { endpointIDHex in
+ do {
+ _ = try judge.judgment()(nil, endpointIDHex)
</file context>
There was a problem hiding this comment.
1 issue found across 2 files (changes from recent commits).
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxJSONCache.swift">
<violation number="1" location="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxJSONCache.swift:176">
P2: Release construction now preserves legacy cache files without performing the identity-validated migration promised by this comment. Wire the migration and cleanup into the factory, or remove the legacy files only after a validated replacement, so stale account state cannot persist and be reused by an unscoped broker.</violation>
</file>
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
| guard let scope else { return file } | ||
| // A legacy snapshot is not account/backend scoped. Never import it | ||
| // into the scoped keychain item, even when its shape happens to decode. | ||
| // Preserve the old copy until an explicit, identity-validated |
There was a problem hiding this comment.
P2: Release construction now preserves legacy cache files without performing the identity-validated migration promised by this comment. Wire the migration and cleanup into the factory, or remove the legacy files only after a validated replacement, so stale account state cannot persist and be reused by an unscoped broker.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxJSONCache.swift, line 176:
<comment>Release construction now preserves legacy cache files without performing the identity-validated migration promised by this comment. Wire the migration and cleanup into the factory, or remove the legacy files only after a validated replacement, so stale account state cannot persist and be reused by an unscoped broker.</comment>
<file context>
@@ -173,9 +173,9 @@ enum IrxBrokerCacheFactory {
- // Remove the old copy so a later unscoped bootstrap cannot resurrect
- // credentials after sign-out or account switching.
- file.clear()
+ // Preserve the old copy until an explicit, identity-validated
+ // migration can replace it; construction must never destroy the only
+ // offline binding or credential state during an upgrade.
</file context>
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
There was a problem hiding this comment.
2 issues found across 3 files (changes from recent commits).
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift">
<violation number="1" location="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift:314">
P1: When `kick()` or `stop()` runs while a frame handler is suspended, the old `route` continuation still applies that frame after the new generation starts. Recheck the generation after each awaited handler and before acknowledging, and bind acknowledgements to the frame's socket so stale frames cannot mutate the new control-plane state.</violation>
</file>
<file name="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxBrokerService.swift">
<violation number="1" location="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxBrokerService.swift:514">
P2: When a cached pushed snapshot contains duplicate relay URLs, `Dictionary(uniqueKeysWithValues:)` traps before the next pass is processed. Build the map with a uniquing closure so malformed or legacy cache data is coalesced safely.</violation>
</file>
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
| case .data(let raw): data = raw | ||
| @unknown default: continue | ||
| } | ||
| await route(data, generation: generation) |
There was a problem hiding this comment.
P1: When kick() or stop() runs while a frame handler is suspended, the old route continuation still applies that frame after the new generation starts. Recheck the generation after each awaited handler and before acknowledging, and bind acknowledgements to the frame's socket so stale frames cannot mutate the new control-plane state.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift, line 314:
<comment>When `kick()` or `stop()` runs while a frame handler is suspended, the old `route` continuation still applies that frame after the new generation starts. Recheck the generation after each awaited handler and before acknowledging, and bind acknowledgements to the frame's socket so stale frames cannot mutate the new control-plane state.</comment>
<file context>
@@ -290,19 +302,21 @@ public actor IrxControlPlaneClient {
@unknown default: continue
}
- await route(data)
+ await route(data, generation: generation)
}
}
</file context>
| var mergedByRelay = Dictionary( | ||
| uniqueKeysWithValues: (cachedSnapshot?.credentials ?? []).map { | ||
| ($0.relayURL, $0) | ||
| }) |
There was a problem hiding this comment.
P2: When a cached pushed snapshot contains duplicate relay URLs, Dictionary(uniqueKeysWithValues:) traps before the next pass is processed. Build the map with a uniquing closure so malformed or legacy cache data is coalesced safely.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxBrokerService.swift, line 514:
<comment>When a cached pushed snapshot contains duplicate relay URLs, `Dictionary(uniqueKeysWithValues:)` traps before the next pass is processed. Build the map with a uniquing closure so malformed or legacy cache data is coalesced safely.</comment>
<file context>
@@ -506,28 +508,44 @@ public actor IrxBrokerService {
+ let cachedSnapshot = credentialCache.load().flatMap { snapshot in
+ snapshot.endpointIDHex == identity.endpointIDHex ? snapshot : nil
+ }
+ var mergedByRelay = Dictionary(
+ uniqueKeysWithValues: (cachedSnapshot?.credentials ?? []).map {
+ ($0.relayURL, $0)
</file context>
| var mergedByRelay = Dictionary( | |
| uniqueKeysWithValues: (cachedSnapshot?.credentials ?? []).map { | |
| ($0.relayURL, $0) | |
| }) | |
| var mergedByRelay = Dictionary( | |
| (cachedSnapshot?.credentials ?? []).map { | |
| ($0.relayURL, $0) | |
| }, | |
| uniquingKeysWith: { _, latest in latest }) |
There was a problem hiding this comment.
3 issues found across 5 files (changes from recent commits).
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift">
<violation number="1" location="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift:431">
P1: When a directory callback fails during an initial snapshot, this line suppresses the ack but `snapshot_complete` still persists that revision as `haveRev`. The next reconnect takes the DO's resumed fast path and skips the failed directory, leaving list-auth state stale; persist the cursor only through the highest contiguous applied revision, or otherwise force an unacknowledged snapshot to replay.</violation>
</file>
<file name="Sources/Mobile/MobileHostIrxRuntime.swift">
<violation number="1" location="Sources/Mobile/MobileHostIrxRuntime.swift:470">
P1: When the secure-store write fails, this guard drops the new directory before enforcing it, so a revoked or delisted peer remains admitted under the old lease. Apply and enforce the snapshot in memory (or clear admission fail-closed) even when persistence fails, while returning `false` so the server retries instead of acknowledging it.</violation>
</file>
<file name="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxBrokerService.swift">
<violation number="1" location="Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxBrokerService.swift:521">
P3: When the new guard rejects pushed credentials because they are already expired (or have refreshAfter >= expiresAt), the journal records reason "stale" in the `!accepted.isEmpty` fallback, the same label already used for monotonic-freshness drops. That conflates the two rejection modes and will mislead anyone debugging the replay-rejection path. Record a distinct reason (e.g. "expired") for the new guard's rejections.</violation>
</file>
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
| } | ||
| // Directory delivery is itself a revisioned fact. Ack only | ||
| // after every consumer reports durable application. | ||
| if applied { await acknowledge(rev: listFact.rev) } |
There was a problem hiding this comment.
P1: When a directory callback fails during an initial snapshot, this line suppresses the ack but snapshot_complete still persists that revision as haveRev. The next reconnect takes the DO's resumed fast path and skips the failed directory, leaving list-auth state stale; persist the cursor only through the highest contiguous applied revision, or otherwise force an unacknowledged snapshot to replay.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/ControlPlane/IrxControlPlaneClient.swift, line 431:
<comment>When a directory callback fails during an initial snapshot, this line suppresses the ack but `snapshot_complete` still persists that revision as `haveRev`. The next reconnect takes the DO's resumed fast path and skips the failed directory, leaving list-auth state stale; persist the cursor only through the highest contiguous applied revision, or otherwise force an unacknowledged snapshot to replay.</comment>
<file context>
@@ -410,26 +411,24 @@ public actor IrxControlPlaneClient {
- await acknowledge(rev: listFact.rev)
+ // Directory delivery is itself a revisioned fact. Ack only
+ // after every consumer reports durable application.
+ if applied { await acknowledge(rev: listFact.rev) }
case "current":
// Explicit freshness re-stamp for the device-list lease.
</file context>
| receivedAtWall: Date(), | ||
| receivedAtMonotonic: .now | ||
| ) | ||
| guard await deviceListStore.persist(snapshot) else { return false } |
There was a problem hiding this comment.
P1: When the secure-store write fails, this guard drops the new directory before enforcing it, so a revoked or delisted peer remains admitted under the old lease. Apply and enforce the snapshot in memory (or clear admission fail-closed) even when persistence fails, while returning false so the server retries instead of acknowledging it.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Sources/Mobile/MobileHostIrxRuntime.swift, line 470:
<comment>When the secure-store write fails, this guard drops the new directory before enforcing it, so a revoked or delisted peer remains admitted under the old lease. Apply and enforce the snapshot in memory (or clear admission fail-closed) even when persistence fails, while returning `false` so the server retries instead of acknowledging it.</comment>
<file context>
@@ -452,23 +453,22 @@ final class MobileHostIrxRuntime {
receivedAtWall: Date(),
receivedAtMonotonic: .now
)
+ guard await deviceListStore.persist(snapshot) else { return false }
deviceListBox.replace(snapshot)
- await deviceListStore.persist(snapshot)
</file context>
| var accepted: [IrxRelayCredential] = [] | ||
| let now = Date() | ||
| for credential in pushed { | ||
| guard credential.expiresAt > now, |
There was a problem hiding this comment.
P3: When the new guard rejects pushed credentials because they are already expired (or have refreshAfter >= expiresAt), the journal records reason "stale" in the !accepted.isEmpty fallback, the same label already used for monotonic-freshness drops. That conflates the two rejection modes and will mislead anyone debugging the replay-rejection path. Record a distinct reason (e.g. "expired") for the new guard's rejections.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At Packages/Shared/CmuxIrxTransport/Sources/CmuxIrxTransport/IrxBrokerService.swift, line 521:
<comment>When the new guard rejects pushed credentials because they are already expired (or have refreshAfter >= expiresAt), the journal records reason "stale" in the `!accepted.isEmpty` fallback, the same label already used for monotonic-freshness drops. That conflates the two rejection modes and will mislead anyone debugging the replay-rejection path. Record a distinct reason (e.g. "expired") for the new guard's rejections.</comment>
<file context>
@@ -516,7 +516,11 @@ public actor IrxBrokerService {
var accepted: [IrxRelayCredential] = []
+ let now = Date()
for credential in pushed {
+ guard credential.expiresAt > now,
+ credential.refreshAfter < credential.expiresAt
+ else { continue }
</file context>
Implements the connection-architecture redesign specified in the living design artifact (hq cmux-assets/irx-client-resilience/irx-sequence-diagram): device-list authorization replaces pair-grant JWS verification, and the account Durable Object becomes the control plane for registration confirmation, device-list sync, acks, and revocation.
Backend (workers/presence): listv2 overlay in the AccountControlPlane DO (status enum + revoked + appVersion/track/capabilities per device, seeded-on-first-sight, confirm-on-hello), directory freshness lease (issuedAt + 24h ttl, snapshot_complete re-stamps), per-device ack(rev) with alarm retries (5s/30s/2m/10m/hourly), POST /v1/control/devices/revoke (instant broadcast, socket close 1008, mint denial). Schemas extended additively + quicktype regenerated (TS + Swift), drift guard green. 240 worker tests.
Clients (both platforms): grantless hello v2; Mac admission = IrxListJudge (EndpointID in list AND !revoked AND lease fresh, synchronous O(1), fail closed on stale/absent); revocation closes live sessions; iOS dial gate + seeded-Mac warning row (en+ja); device-list lease persisted in Keychain (ThisDeviceOnly, keyed account|backend); broker JSON caches move to Keychain in Release with one-way file migration; control-plane client sends client info + acks and starts before broker calls. 62 + 338 package tests.
Transport hardening found by the gates: keepalive now requires two consecutive pong misses before declaring death (a lone 2s stall severed an otherwise-perfect 9-min session; the Mac's pong landed 85ms past the deadline), and automatic path mode now actually authorizes NAT traversal after admission (deferNatTraversalUntilAuthorized was set at bind but nothing ever authorized, so sessions could never leave the relay — the known "direct-path unverified" v1 gap).
Also fixes a main compile break: AgentHibernationRecord.processLiveness had an inline default, excluding it from the memberwise init used by +Records.swift (#10658 fallout; the PR lane has no macOS compile gate).
Acceptance gates run live (isolated sim as iPhone, tagged Mac, isolated dev worker cmux-presence-dev-listauth proxying staging; harness scripts/listauth-gates.py; evidence in hq cmux-assets/irx-client-resilience/listauth-gates-evidence/):
Deferred by design decision: the full sign-in→terminal XCUITest (post-implementation design pass), hint-rev acks, Mac account-erase seam for broker caches (pre-existing), MobileMacListAuthState singleton seam.
🤖 Generated with Claude Code
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.Summary by cubic
Replaces pair-grant JWS verification with device-list authorization and moves directory sync, acks, and revocation into the per-account control-plane Durable Object. Admission checks the Mac's endpoint against a persisted 24h-leased device list; revoked devices are cut immediately and denied future admission, and clients ack applied control-plane revisions so the server's retry ladder stops at them.
Transport and client behavior
Control plane and infrastructure
schemas/control-plane/with a CI drift guard; transport state is namespaced per bundle and broker so dev and production builds can't overwrite each other's trust keys.AgentHibernationRecord.processLivenessto the memberwise initializer, fixing a main compile break.Written for commit dfe4185. Summary will update on new commits.
Summary by CodeRabbit