Repository navigation
irx: never stop attempting; cap every backoff at 30 minutes - #15450
azooz2003-bit wants to merge 3 commits into
Conversation
Yesterday cmux NIGHTLY stopped renewing relay credentials at 03:00Z and stayed unreachable from iOS for 17 hours without one log line saying why. Every path that can stop renewals was unlogged; the 12h unified-log retention then erased the failure window. This adds journal events at each silent exit so the next wedge is attributable from retained logs: - V2ControlService: session-ready, socket-failed, socket-open-failed, http-mode-entered, run-backing-off, run-stopped-terminal, maintenance-scheduled/-not-scheduled/-planned (with per-schema due times and cooldown deferrals)/-exited (with reason), refresh-succeeded and refresh-failed per schema, cooldown-set with source and delay, and persist-failed, all journaled from the service so a stalled snapshot consumer cannot hide them. - MobileHostIrxRuntime: credentials-received, endpoint-ready-skipped with reason, and a renewal watchdog that reads the service directly every 5 minutes and journals credential-renewal-overdue and snapshot-apply-stalled (apply completion tracked via defer so a hang inside apply stays visible). - IrxEndpoint/installer: relay-rotation-skipped/-deferred, relay-credential-install-started/-superseded, and relay-credential-unusable, making a hung native install visible as a started event with no outcome. The journal is injected through V2ControlDependencies (defaulted nil) and wired on both macOS and iOS. Events carry schema names, failure codes, counts, and durations only, never tokens. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The 2026-09-28 NIGHTLY wedge showed a Mac silently stop renewing relay credentials for 17 hours. Whatever the exact trigger was, the code had four ways to stop attempting entirely, and none of them is acceptable: a reachable Mac must keep announcing itself at a bounded cadence. - terminal() now stops the run only for a deliberate revocation of this device's authority (server code or 1008 close reason). scopeMismatch, persistenceFailed, capacityExceeded, invalidWireData, non-revocation policy closes and 4xx responses all retry: one malformed frame or one failed disk write used to kill connectivity until app relaunch. - V2ControlService.maximumBackoff (30 min) caps every delay: the client_upgrade_required ladder drops from 1h/6h/24h to 10m/30m/30m, rate-limited and HTTP 429 cooldowns honor the server retry-after only up to the ceiling, and retryDelay clamps whatever an accumulated cooldown says. Relay renewal retry stays ~30s per failure, far under the relay.request rate limit. - The renewal loop no longer exits when it wakes during the brief socket-gone window; it idles at 30s (journaled as maintenance-idle) until the transport returns or the reconnect owner changes status. - MobileHostIrxRuntime.activationRetryDelay: the server retry-after floor is capped at 30 minutes; the ladder already capped at 5. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Important Review skippedWe couldn't safely recover the incremental review. No full review was started, and the last reviewed checkpoint was preserved. Retry later, or explicitly request a full review by commenting You can disable this status message by setting the Use the checkbox below for a quick retry:
📝 WalkthroughWalkthroughThe changes add structured IRX lifecycle events, revise terminal and retry handling, cap retry and cooldown delays, and keep maintenance running when transport is unavailable. The mobile runtime adds a renewal watchdog and logs endpoint credential readiness and rotation events. ChangesIRX renewal and retry reliability
Estimated code review effort: 3 (Moderate) | ~25 minutes Change: Bug fix Sequence Diagram(s)sequenceDiagram
participant MobileHostIrxRuntime
participant V2ControlService
participant IrxJournal
MobileHostIrxRuntime->>V2ControlService: Provision with journal dependency
V2ControlService->>IrxJournal: Record control and refresh events
MobileHostIrxRuntime->>V2ControlService: Read snapshot during five-minute watchdog check
MobileHostIrxRuntime->>IrxJournal: Record overdue renewal and unapplied snapshot events
Suggested reviewers: ✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Dogfood tours of
|
|
Automatic catch-up couldn't merge Label |
Why
Directive from the 2026-09-28 NIGHTLY wedge (17h of silent relay-credential non-renewal, #15443): there must be no path that stops relay credential refreshing or connection attempts, and no backoff anywhere may exceed 30 minutes. The V2 control service had four stop-entirely classes:
scopeMismatch,persistenceFailed(single failed disk write),capacityExceeded,invalidWireData(one malformed server frame), or any 1008/1009 close setstatus = .stopped, cleared the run, and nothing short of app relaunch or a manual Settings retry recovered. Even wake/foreground no-ops on a stopped run.client_upgrade_requireddeferred a schema 1h → 6h → 24h;rate_limitedand HTTP 429 honored uncapped server retry-afters;retryDelayinherited all of those throughmax()..ready.activationRetryDelay's ladder caps at 5 minutes, but an uncapped server retry-after floor could override it to hours.What
terminal()stops the run only for deliberate revocation of this device's authority: serverdevice_revoked/team_access_revoked, or a 1008 close carrying one of those reasons. Everything else retries on the reconnect ladder.V2ControlService.maximumBackoff = 30mincapsretryDelay, the upgrade-required ladder (now 10m/30m/30m), rate-limited cooldowns, and HTTP 429 cooldowns. Relay refresh failures keep their ~30s maintenance retry, far under the server'srelay.requestper-hour budget.maintenance-idle) instead of exiting when the transport is momentarily gone.activationRetryDelayclamps the retry-after floor at 30 minutes.Rate-limit safety: worst case per schema stays ≥30s between attempts (maintenance failure cooldown), and reconnects keep exponential growth up to the ceiling, so a genuinely broken client asks the server at most ~2/min transiently and settles to ≥30s cadence, orders of magnitude under
relay.requestperHour(1200).Testing
swift testinPackages/Shared/CmuxIrxTransport: 217 tests green. Spec-change tests updated and extended:onlyRevocationStopsEverythingElseRetriesWithinTheBackoffCeiling(revocation reasons stop; policy/wire/persistence/4xx do not; 6h retry-after and 24h accumulated cooldown both clamp to ≤1800s), retired-schema cooldown now asserts the 300-1800s bound, journalcooldown-setdelay bound follows.Stacked on #15443 (journal events used by the new idle path). Sibling: #15446 ships these events to Axiom, so
maintenance-idleloops and capped cooldowns are visible in the sink.Changelog
Fixed: a Mac could permanently stop renewing relay credentials or reconnecting after a transient failure; all connection and credential retries now continue at a bounded cadence and no backoff exceeds 30 minutes.
🤖 Generated with Claude Code
Summary by cubic
Prevents relay-credential renewal and reconnection from ever stopping permanently or backing off for more than 30 minutes, so a wedged Mac keeps announcing itself instead of going silent for hours.
terminal()now stops the run only for a deliberate revocation (device_revoked/team_access_revoked); wire-shape, persistence, policy, and 4xx failures all retry.Written for commit a876319. Summary will update on new commits.
Summary by CodeRabbit