fix(iroh-v2): check the team socket cap before opening the session - #15340
teamleaderleo merged 2 commits into
Conversation
… work
A team at TEAM_SOCKET_LIMIT still runs the whole open path for every further
/socket attempt, because the cap is checked after broker.open has returned.
This adds a runtime case that pins the two side effects that must not happen
on a refused connection: the team revision must not move, and no
directory.changed.v1 frame may be pushed to the sockets already connected.
The cap becomes a single overridable value so both admission paths read it,
which also lets the runtime fixture lower it. 4096 live sockets are not
reachable under Miniflare, and the dashboard path had the number copied by
hand, so the two caps could drift apart without anyone noticing.
Red on this commit:
461 | expect(await control.teamRevision()).toBe(before);
^
error: expect(received).toBe(expected)
Expected: 3
Received: 4
(fail) a team at its socket cap sheds the next socket before it mutates team state
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
TEAM_SOCKET_LIMIT was checked after broker.open had already returned, so a team sitting at the cap did the entire open path for every further /socket attempt and then threw the work away: an Ed25519 verification over caller-supplied bytes, a device-proof write, possibly a challenge row, and recording fresh authority, which bumps the team revision. The revision bump is then fanned out as directory.changed.v1 to every socket in the object and to every dashboard socket. A cap whose job is to shed load was instead multiplying it, and the clients already connected paid for the attempts that got rejected. Move the upgrade and cap checks ahead of broker.open, scoped to /socket so /session admission is unchanged. Neither check reads anything open() produces, and the failure telemetry keeps reporting stage "accept" so existing queries on route=socket still work. Green on this commit, same command as the previous one: 13 pass 0 fail 100 expect() calls Ran 13 tests across 1 file. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Warning Review limit reachedNext included review available in 1 minute. View limit detailsLimit details: You’ve used all 10 included reviews currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Repository: manaflow-ai/cmux/.coderabbit.yaml Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (4)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
All contributors have signed the CLA ✍️ ✅ |
|
Cross-model review (Codex gpt-5.6-sol)
|
|
Review: a review subagent went through this at The question that mattered: does moving the cap check ahead of Independently verified rather than taken on trust: the red/green receipt is exact. On the parent commit the test fails at Left, all nits, none needing a code change:
Licensing: the diff stays inside Verification the reviewer ran on |
|
Merging on green under the standing rule for fix PRs (skip team review, dogfood, merge on green). No fleet dogfood evidence on this one, and the reason is structural rather than a skipped step: #8029 turned off Vercel branch previews while keeping |
|
Merge receipt for |
0e298fb ci: wait for the product's canonical root instead of compiling beside it (manaflow-ai#15379) 3088273 ci: UI test runs adopt compile admission's product, skip the re-upload, and report progress (manaflow-ai#15331) b681e7e Keep a pending banner quiet once its pane is focused (manaflow-ai#15357) 03a2f6e Record that cloud_vm_sessions.attachment_count is cumulative (manaflow-ai#15321) 48258b4 fix(iroh-v2): check the team socket cap before opening the session (manaflow-ai#15340) 2638d56 Agent activity reorder follow-ups: group on-top check, search, subtitle (manaflow-ai#15362) 9ed83fd Dogfood journey: record whether a paused Cloud machine is asleep (manaflow-ai#15293) 7171ea8 Add app.tabBarVisibility to hide the pane tab bar when a pane has one tab (manaflow-ai#15294) 8743ec8 test: stop Computer Use onboarding tests waiting out the helper status deadline (manaflow-ai#15329) 6e4f1da ci: drain the snapshot's owned queue by what the machines finished since (manaflow-ai#15374) 9373164 ci: queue a pull request's admission for a root runner when Blacksmith's wait is longer (manaflow-ai#15376) 634a155 test: expect injected pane attention accent (manaflow-ai#15370) cd030e9 Keep a named Cloud machine's prompt name instead of flipping to its slug (manaflow-ai#15288) 24ee0ee Exit 1 when cmux terminal screen wait times out (manaflow-ai#15282) 1b857ac test: cover a live Codex turn owner keeping its turn on SessionStart (manaflow-ai#13588) 56ec600 PR media: prune media of long-closed pull requests (manaflow-ai#15364) 4898cde ci: bound the SwiftPM scratch holder and cache scratch sizes (manaflow-ai#15366) # Conflicts: # .github/workflows/ci-guards.yml # .github/workflows/ci.yml # .github/workflows/test-e2e.yml
The failure
TeamControl.fetchcheckedTEAM_SOCKET_LIMITafterawait broker.open(...)hadalready returned, and after
scheduleChangeshad already queued the broadcast.So for a team sitting at 4096 sockets, every further
/socketattempt ran thewhole open path and then threw the result away:
verifyDeviceSignature, an Ed25519 verification over caller-supplied bytesconsumeDeviceProof, a storage writeissueChallengefor a device that is not enrolled, another storage writeobserveAuthority, which records fresh authority and increments the teamrevision
scheduleChanges, whichwaitUntils adirectory.changed.v1frame to everysocket in the object plus a
dashboard.broadcastto every dashboard socketand only then returned 429.
The cost does not stop at the rejected caller. Each refused attempt woke every
client already connected to that team with a revision invalidation, and each of
those clients answers a revision invalidation by re-requesting the directory.
A cap that exists to shed load was amplifying it, worst exactly when the team
was already at capacity.
Why this is a defect and not a design choice
The function's own
stagetracking says the order was meant to be the other wayaround:
stage = "open"is set beforebroker.open,stage = "accept"afterit, and the cap check sat in the
"accept"stage. Admission was intended toprecede the session open.
Neither relocated check reads anything
broker.openproduces. The cap readsthis.ctx.getWebSockets().length; the upgrade check reads a request header.resultis used only from the line after them onwards. The sibling dashboardpath in
dashboard-control.tsalready checks its cap before it reserves oraccepts anything, so the native path was the outlier.
The rejected caller sees exactly what it saw before: 429, retryable, 5000 ms
backoff. Failure telemetry still reports
route: "socket", stage: "accept", soexisting queries keep working.
Changes
src/team-control.ts: the upgrade and cap checks move ahead ofbroker.open, scoped toincoming.path === "/socket"so/sessionadmission is untouched.
src/team-control.tsandsrc/dashboard-control.ts: the cap becomes oneoverridable value that both admission paths read.
dashboard-control.tshadthe literal
4096copied by hand, so the two caps could drift apart silently.It is passed through the existing
Serviceshooks rather than imported,because importing it would make the module graph cyclic.
e2e/control-runtime.test.tsande2e/control-worker.ts: a runtime case thatholds a socket open, lowers the cap through a test-only fixture override, and
pins both side effects that must not happen on a refused connection. The
fixture lowers the cap because 4096 live sockets are not reachable under
Miniflare; production never subclasses
TeamControl, so it always reads the4096 constant.
Red and green
Same command both times:
bun test e2e/control-runtime.test.tsinworkers/iroh-v2.Red, on the test commit (
7ad6689):The second assertion is red on that commit too. With the revision assertion
removed so execution reaches it, the frame that a rejected connection pushed to
the socket already connected is:
Green, on the fix commit (
28341dd):Also green on the fix commit:
bun run checkinworkers/iroh-v2(boundary:check, contracts:check,types:check, typecheck, test):
79 pass, 0 fail, Ran 79 tests across 17 filesbash scripts/test-runtime.sh, all three Miniflare suites:13 pass,8 pass,22 pass, 0 failNo macOS or Xcode build was run; this is a Cloudflare Worker only, and nothing
here ships in the macOS or iOS app.
Changelog
Fixed: a cmux team already at its control-socket limit no longer makes every
connected client resynchronize its device directory each time a further
connection is refused.
🤖 Generated with Claude Code
Summary by cubic
Fixes the team socket cap so refused connections no longer run the session-open path or wake every connected client.
Previously
TeamControl.fetchcheckedTEAM_SOCKET_LIMITafterbroker.openhad already returned andscheduleChangeshad queued a broadcast. A team at 4096 sockets ran the whole open path — signature verification, storage writes, authority observation, team revision bump — for every rejected attempt, then pushed that revision invalidation to every socket already connected, and each client re-requested the directory in response. The load-shedding cap was amplifying load.broker.open, scoped to/socketso/sessionadmission is untouched. Failed requests still return 429 with the same telemetry stage.team-control.tsanddashboard-control.tsas one overridable value supplied through the existingServiceshooks, since the dashboard had the literal4096copied by hand.directory.changed.v1frame may reach connected sockets. The test fixture lowers the cap since 4096 live sockets aren't reachable under Miniflare.Written for commit 28341dd. Summary will update on new commits.