fix(desktop): reclaim idle pool backend for foreground dials (#102281 follow-up) - #104871
bounce12340 wants to merge 3 commits into
Conversation
Co-authored-by: Josh Tsai <bounce12340@users.noreply.github.com>
Co-authored-by: Josh Tsai <bounce12340@users.noreply.github.com>
Co-authored-by: Josh Tsai <bounce12340@users.noreply.github.com>
SummaryFollow-up to #102281: a foreground dial may now reclaim a keepalive-fresh but idle pool backend instead of stalling on a full pool. Freshness alone no longer implies busy — the renderer publishes an explicit Findings
VerdictLooks good. The explicit lease cleanly separates "interested" from "busy", with the unknown state defaulting to protected. Safe to merge. |
|
Related / complementary work: #105390 ( This PR (#104871) is the other half of the #102281 follow-up: when the pool is already full of keepalive-fresh idle residents, reclaim the oldest idle ( The other author notes no line-level overlap with this change. Happy for maintainers to review both together — they can coexist, or we can follow guidance on how to split/merge. |
|
Superseded by #113280. Your trigger point (a foreground dial at |
… an admission fence A foreground bot open against a full 3-slot local pool waited 30s for a slot and timed out (production: 7,122 slot waits, 6,305 timeouts, ~790 cancelled starts in 10 days on a 17-profile host). Nothing could free a slot: a child's hard lease lives until the child exits (main.ts child 'exit' handler, teardownFailedLocalBackend, stopPoolBackend after the bounded SIGTERM->SIGKILL exit), while LRU eviction and the idle reaper key off lastActiveAt, which the renderer refreshes every 60s for every open socket. A bot-tile-pinned resident is keepalive-fresh forever. Occupied is not busy. New `electron/pool-retire.ts` (pure, DI-testable like pool-spawn-coordinator): a foreground dial that finds `activeCount >= poolMaxBackends()` may retire ONE resident, under these rules: * Proof is the backend's. Each LRU candidate is probed over `/api/health/idle` (running sessions + running cron jobs + prompts waiting on a human). Only `true` is idle; `false` and `null` (older runtime 404, probe error, unreadable ledger) are busy. The renderer-published `activeTurn` lease is an early skip, never the proof; entries no longer initialise it to false, so a fresh backend is not evictable by default. * Admission fence: concurrent foreground dials share one retirement and one probe; the coordinator hands the freed slot to the ticket that queued first. * Identity recheck (`pool.get(key) === entry`, no new turn lease) after every await, and a re-probe immediately before the stop. * The waiter's `coordinator.request()` is issued BEFORE the stop so it sits at the queue head; the slot is released only by the retired child's real exit through the existing stopPoolBackend -> releaseLocalBackendSlot path. * `hermes:pool:retiring` is broadcast to every window before the SIGTERM so the renderer parks the scope instead of redialing into the vacated slot. main.ts grows only the trigger point, the HTTP probe, the broadcast and the retirer construction. Kept from #104871 with credit: the trigger point before `coordinator.request`, the `touchBackend(profile, options)` IPC widening (preload.ts / global.d.ts), and the LRU-among-eligible selector shape. Tests (pool-retire.test.ts): fence with two concurrent tickets over a real LocalBackendSpawnCoordinator (one probe, one SIGTERM, slot granted only after the simulated exit, exactly one waiter served); identity recheck aborts on a swapped entry or a lease published after the probe; idle null / cron-running ineligible with fall-through to the queue; renderer activeTurn:false loses to a backend re-probe. Co-authored-by: bounce12340 <128559392+bounce12340@users.noreply.github.com>
When main retires a pooled backend for a foreground open, the renderer entry
riding it is still wantOpen: the socket's 'closed' state ran
scheduleReconnect -> reconnectSecondary -> openSecondary('background') and
immediately re-queued a spawn for the slot the retirement had just freed. The
wake/focus nudge (reconnectSecondaryGateways) re-arms parked entries too, so
even a stall-parked scope would have redialed on the next window focus.
Listen for `hermes:pool:retiring` in use-gateway-boot and route it through a
new `parkSecondariesForRetiredBackend(poolKey)`: every scope riding that
child (bare profile or `conn:local::<profile>`; both ride one local child)
gets wantOpen=false, its reconnect timer cleared, and a sticky `retiredByPool`
flag that the nudge skips. The entry stays, so bot tiles keep their card
without a live socket. Only an explicit open of the scope (rearmSecondary,
reached from ensureGatewayForAgent / openGatewayForAgent /
ensureGatewayForProfile / ensureActiveGatewayOpen) clears the flag and dials
again.
Turn-lease reporting (shape from #104871): retainGatewayForSessionTurn and
touchSecondaryGateways publish `activeTurn` per scope through the widened
touchBackend IPC. On release the SCOPE's lease state is reported, not the
single lease's, so a second session on the same backend keeps it marked.
Test (gateway-connection-lifecycle.test.ts): a retired local scope does not
redial from the nudge across ~400s of fake time while an unrelated remote
scope is untouched; an explicit open re-arms it and the nudge treats it
normally afterwards. RED on base: the nudge redialed the retired scope.
Co-authored-by: bounce12340 <128559392+bounce12340@users.noreply.github.com>
… an admission fence A foreground bot open against a full 3-slot local pool waited 30s for a slot and timed out (production: 7,122 slot waits, 6,305 timeouts, ~790 cancelled starts in 10 days on a 17-profile host). Nothing could free a slot: a child's hard lease lives until the child exits (main.ts child 'exit' handler, teardownFailedLocalBackend, stopPoolBackend after the bounded SIGTERM->SIGKILL exit), while LRU eviction and the idle reaper key off lastActiveAt, which the renderer refreshes every 60s for every open socket. A bot-tile-pinned resident is keepalive-fresh forever. Occupied is not busy. New `electron/pool-retire.ts` (pure, DI-testable like pool-spawn-coordinator): a foreground dial that finds `activeCount >= poolMaxBackends()` may retire ONE resident, under these rules: * Proof is the backend's. Each LRU candidate is probed over `/api/health/idle` (running sessions + running cron jobs + prompts waiting on a human). Only `true` is idle; `false` and `null` (older runtime 404, probe error, unreadable ledger) are busy. The renderer-published `activeTurn` lease is an early skip, never the proof; entries no longer initialise it to false, so a fresh backend is not evictable by default. * Admission fence: concurrent foreground dials share one retirement and one probe; the coordinator hands the freed slot to the ticket that queued first. * Identity recheck (`pool.get(key) === entry`, no new turn lease) after every await, and a re-probe immediately before the stop. * The waiter's `coordinator.request()` is issued BEFORE the stop so it sits at the queue head; the slot is released only by the retired child's real exit through the existing stopPoolBackend -> releaseLocalBackendSlot path. * `hermes:pool:retiring` is broadcast to every window before the SIGTERM so the renderer parks the scope instead of redialing into the vacated slot. main.ts grows only the trigger point, the HTTP probe, the broadcast and the retirer construction. Kept from #104871 with credit: the trigger point before `coordinator.request`, the `touchBackend(profile, options)` IPC widening (preload.ts / global.d.ts), and the LRU-among-eligible selector shape. Tests (pool-retire.test.ts): fence with two concurrent tickets over a real LocalBackendSpawnCoordinator (one probe, one SIGTERM, slot granted only after the simulated exit, exactly one waiter served); identity recheck aborts on a swapped entry or a lease published after the probe; idle null / cron-running ineligible with fall-through to the queue; renderer activeTurn:false loses to a backend re-probe. Co-authored-by: bounce12340 <128559392+bounce12340@users.noreply.github.com>
When main retires a pooled backend for a foreground open, the renderer entry
riding it is still wantOpen: the socket's 'closed' state ran
scheduleReconnect -> reconnectSecondary -> openSecondary('background') and
immediately re-queued a spawn for the slot the retirement had just freed. The
wake/focus nudge (reconnectSecondaryGateways) re-arms parked entries too, so
even a stall-parked scope would have redialed on the next window focus.
Listen for `hermes:pool:retiring` in use-gateway-boot and route it through a
new `parkSecondariesForRetiredBackend(poolKey)`: every scope riding that
child (bare profile or `conn:local::<profile>`; both ride one local child)
gets wantOpen=false, its reconnect timer cleared, and a sticky `retiredByPool`
flag that the nudge skips. The entry stays, so bot tiles keep their card
without a live socket. Only an explicit open of the scope (rearmSecondary,
reached from ensureGatewayForAgent / openGatewayForAgent /
ensureGatewayForProfile / ensureActiveGatewayOpen) clears the flag and dials
again.
Turn-lease reporting (shape from #104871): retainGatewayForSessionTurn and
touchSecondaryGateways publish `activeTurn` per scope through the widened
touchBackend IPC. On release the SCOPE's lease state is reported, not the
single lease's, so a second session on the same backend keeps it marked.
Test (gateway-connection-lifecycle.test.ts): a retired local scope does not
redial from the nudge across ~400s of fake time while an unrelated remote
scope is untouched; an explicit open re-arms it and the nudge treats it
normally afterwards. RED on base: the nudge redialed the retired scope.
Co-authored-by: bounce12340 <128559392+bounce12340@users.noreply.github.com>
Summary
activeTurn === falselease; unknown activity fails closed; active turns are never reclaimed.touchBackend(..., { activeTurn }).Related: #102281 (follow-up to #104139; issue remains closed — still repro after merge).
Test plan
npx vitest run --project electron electron/pool-eviction.test.ts(10 passed per draft)maxBackends=3, fill idle residents then open another bot — oldest idle reclaimed