fix(network): bound direct-path response-start timeout and retry on fresh socket (#10214) - #10528
Conversation
8a90142 to
824eeba
Compare
|
Não foi possível revisar o código com segurança: /tmp/review-prs-0816-diffs/10528.diff está vazio (0 bytes). Além disso, o PR aberto aponta para main, enquanto a base solicitada é release/v3.8.50. A metadata disponível descreve 584.103 adições, 105.697 remoções e 100 arquivos, portanto não é seguro inferir o patch pretendido a partir dela. Solicito um diff não vazio contra release/v3.8.50 e o retarget do PR antes de avaliar feature, segurança, testes, impacto cross-layer ou a regra #18. O alerta base-red #9985 permanece herdado e não foi atribuído ao PR. |
|
The PR is now correctly targeted at |
8e2b729 to
af9b9f2
Compare
af9b9f2 to
142ae93
Compare
…dispatcher-timeout-10214 fix(network): bound direct-path response-start timeout and retry on fresh socket (diegosouzapw#10214)
Fixes #10214 — direct (no-proxy) egress stalls opencode-go / command-code until a service restart.
Root cause
The direct egress path in
open-sse/utils/proxyFetch.tsfunnels every request through the cached round-robin pool of 32 one-connection Agents (getDefaultDispatcher()). A keep-alive pooled socket that silently drops (half-open, no RST — VPN reconnects, edge idle-drops) surfaces no transport error, so the existing fresh-socket retry (which only fires onUND_ERR/ECONNRESET/fetch failed) never triggers: the request hangs until undici'sheadersTimeout(600s default) or the caller's per-model deadline, then fails. A service restart drops the stale sockets, which is why the same providers recover immediately — and re-wedge under sustained traffic.Verified live:
model=orchestratorthrough a long-running gateway: opencode-go leg hangs exactly 120s (combo timeout), zero TCP connections from the gateway process during the hang.Fix
Bound each direct attempt's response-start window (fetch promise = headers only; the body streams through untouched afterwards):
OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS(0 disables — previous behavior).getRetryDispatcher) — a brand-new socket that cannot be the zombie — instead of consuming the full 600s headersTimeout/caller deadline.DIRECT_RESPONSE_START_TIMEOUTerror so the combo fails over to the next target; no native-fetch fallback (native fetch pools too and would hang identically).Tests
tests/unit/proxyfetch-direct-response-start-timeout-10214.test.ts(3 tests):getDefaultDispatcher→getRetryDispatcher).DIRECT_RESPONSE_START_TIMEOUTsurfaces, native fallback never fires.Plus all 44 existing proxyFetch/proxyDispatcher tests: 47/47 pass.
End-to-end verification
A/B against a local server that accepts the first POST and never responds (real half-open zombie socket), with
OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS=5000:Log evidence from the fixed bundle:
[ProxyFetch] Direct response-start timeout (5000ms) on pooled dispatcher — retrying on fresh no-keep-alive dispatcher: 127.0.0.1:19090Also verified the fixed bundle handles the real opencode-go path:
model=orchestrator→ 200 in 1.97s streamingdeepseek-v4-flash.