fix(dispatcher): stale keep-alive burst on local egress - #14315
Merged
diegosouzapw merged 2 commits intoSep 29, 2026
Merged
diegosouzapw merged 2 commits into
diegosouzapw merged 2 commits into
Conversation
Docker Desktop NAT silently drops idle keep-alive sockets inside the direct round-robin pool, and getDefaultDispatcher() never reaps them on PROXY_UNREACHABLE — every .internal/.local hostname (host.docker.internal plus all local LLM providers wired through it) can hit a 30s ECONNREFUSED burst after the keepAliveMaxTimeout window expires. Three coordinated changes: 1. Hostname-aware dispatcher options. getDispatcherOptions() now accepts the target hostname and shortens keepAliveMaxTimeout to 1000ms and autoSelectFamilyAttemptTimeout to 200ms when the hostname matches *.internal / *.local. The latter matters because the IPv6 form of host.docker.internal (fdc4:f303:9324::254) is ENETUNREACH inside the container; the default 1s Happy-Eyeballs wait was pure latency on every healthy request. 2. Parallel cache for local-egress dispatchers. Added LOCAL_DEFAULT / LOCAL_RETRY_DISPATCHER_KEY in proxyDispatcherCache and routed getDefaultDispatcher(hostname?) / getRetryDispatcher(hostname?) to them when isLocalEgressHostname(hostname). The cloud-upstream pool keeps its wider keep-alive; the local-egress pool gets the tighter settings without contaminating each other. 3. Cache invalidation on PROXY_UNREACHABLE for local egress. proxyFetch now calls clearDispatcherCache() before _nativeFallback when the target hostname is local-egress and the error is a proxy unreachable, forcing the next request to rebuild the pool with fresh sockets. Backward compatible: every existing caller of getDefaultDispatcher() / getRetryDispatcher() / __getDefaultDispatcherOptionsForTest() continues to work — the hostname parameter is optional and defaults to the cloud behaviour. Cherry-picked onto release/v3.8.51; one conflict in proxyFetch.ts resolved by taking upstream rename of TlsProfileResult to a named type alias (no semantic change).
…apw#14315 isLocalEgressHostname(), the LOCAL_DEFAULT/LOCAL_RETRY dispatcher cache routing, the shortened local-egress timeouts, and the PROXY_UNREACHABLE-on-local-egress clearDispatcherCache() branch in proxyFetch.ts had zero test coverage. Adds a focused suite covering the 4 cases flagged in review: hostname boundary matching, the hostname-branched dispatcher options, LOCAL_* cache-slot routing (vs. the shared DEFAULT/RETRY slots), clearDispatcherCache() clearing both local slots, and that a local-egress PROXY_UNREACHABLE tears down the local dispatcher pool while a cloud-upstream one does not. Also prunes a stale eslint-suppressions.json entry for proxyDispatcher.ts (an unused-vars violation the branch's own diff had already fixed; the frozen count no longer matched reality and blocked lint-staged). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
On macOS Docker Desktop, every
*.internal/*.localhostname thecontainer reaches through
host.docker.internal(every local LLMprovider wired through the proxy: oMLX, lightning-mlx, mtplx, Breeze TTS,
Parakeet STT, Kokoro TTS) can hit a 30-second ECONNREFUSED burst when
Docker Desktop's NAT silently drops idle keep-alive sockets inside the
direct round-robin pool. Triggers: a request lands on a stale pooled
socket, retries, gets a fresh one, hits the same dead path, then
cooldown/lockout kicks in.
Three concrete amplifiers:
proxyDispatcher.ts:73-76setsautoSelectFamilyAttemptTimeout: 1000,but the IPv6 form of
host.docker.internal(
fdc4:f303:9324::254) isENETUNREACHinside the container (theIPv4 form
192.168.65.254is the working address). Every healthyrequest waits 1s on the dead family before getting ECONNREFUSED on v4.
getDefaultDispatcher()caches the round-robin pool inglobalThisand never reaps onPROXY_UNREACHABLE— a single stalesocket poisons the pool for all subsequent requests until process
restart.
keepAliveMaxTimeout(4s) sits squarely inside the NAT'ssilent-drop window for local egress.
Fix
Three coordinated changes, all backward-compatible:
1. Hostname-aware dispatcher options —
getDispatcherOptions(hostname?)shortenskeepAliveMaxTimeoutto 1000ms andautoSelectFamilyAttemptTimeoutto200ms when the hostname matches
*.internal/*.local. Newexported
isLocalEgressHostname()helper.2. Parallel cache for local-egress dispatchers —
proxyDispatcherCache.tsaddsLOCAL_DEFAULT_DISPATCHER_KEY/LOCAL_RETRY_DISPATCHER_KEYwith matching accessors.getDefaultDispatcher(hostname?)/getRetryDispatcher(hostname?)route to them when the hostname is local-egress. The cloud-upstream
pool keeps its wider keep-alive.
3. Cache invalidation on PROXY_UNREACHABLE for local egress —
proxyFetch.tscallsclearDispatcherCache()before_nativeFallbackwhen the error is
PROXY_UNREACHABLEand the target hostname islocal-egress, forcing the next request to rebuild the pool with fresh
sockets.
Files
open-sse/utils/proxyDispatcher.ts(+63 / −8)open-sse/utils/proxyDispatcherCache.ts(+29)open-sse/utils/proxyFetch.ts(+22 / −1)Backward compatibility
Every existing caller of
getDefaultDispatcher()/getRetryDispatcher()/__getDefaultDispatcherOptionsForTest()keepsworking — the
hostname?parameter is optional and defaults to theexisting cloud behaviour. No config / env / DB changes.
Verification
proxyfetch-direct-response-start-timeout-10214direct-dispatcher-pipelining-4580proxy-egress-isolation-bddproxy-fetchproxy-concurrency-keepalive-regressionweb-cookie-validation-proxy-7058socks-connect-timeout-e2ePre-existing failures on the cycle (
uc-videopersona-timeout,models-catalogJina/GLM-5.2,tunnel-routessanitize,zai-streamerror code) reproduce on
mainunchanged — none introduced by this PR.Notes
Cherry-picked onto
release/v3.8.51from a separate working branch.One conflict in
proxyFetch.tsresolved by taking upstream's rename ofthe inline
tlsProfileForProviderreturn type to a namedTlsProfileResultalias (no semantic change).The companion AgentBridge MITM autostart and
call_logs.idUNIQUE racefixes are intentionally not included here — they're independent fixes
and will land separately.
Related work (not duplicates)
These recently merged neighbours touch nearby layers but do not
supersede this PR — included here to head off the "is this duplicate?"
review question:
fix(db): rotate proxy pools on the chat path like the registry does— DB-layerresolveProxyForConnectioncache-freeze fix for the chat-path proxy resolver. Does not touch undici dispatcher cache, keep-alive, or local-egress hostnames.fix(sse): stop direct fetch retry reusing pooled flat response-start budget—directHeadersTimeoutMsper-attempt budget for the response-start timeout. Different layer (response-start timeout, not idle socket lifetime).fix(sse): pin DNS on the three public-only image fetch sites—pinDns: truefor outbound image fetches. Different layer (DNS, not undici dispatcher).fix(proxy): skip bare TCP health probe for SOCKS5 data plane— SOCKS5-only health probe change. No dispatcher cache work.feat(proxy): support multiple local core endpoints, one per line— proxy-tunnel loopback endpoint registration. Different problem surface (proxy pool rotation in the DB layer).feat(sse): per-egress pacing + fleet-wide backoff for opencode rotation (opt-in)— opt-in viaOPENCODE_EGRESS_THROTTLE_ENABLED=1. OpenCode-only IP-bucketed burst throttle, behind a flag. Different layer.None of the above touches the three properties this PR changes: (1) hostname-keyed
getDispatcherOptions, (2) parallelLOCAL_*_DISPATCHER_KEYcache slots, (3)clearDispatcherCache()invocation onPROXY_UNREACHABLEfor local-egress.