fix(sse): stop an empty Claude stream from emptying the whole combo - #14314
Conversation
A client with four Claude accounts kept getting
429 Service temporarily unavailable: all targets were skipped by
pre-dispatch filters
while two of those accounts were healthy and not in cooldown. Clearing the
cooldowns changed nothing, because the cooldowns were never the problem.
Three defects compounded:
1. Anthropic rejects unknown top-level fields outright — `400 safeguards:
Extra inputs are not permitted`. Newer Claude Code builds send
`safeguards`, and the pure-passthrough path forwarded it verbatim, so
every such request 400'd. Strip it, exactly as top_p is already stripped.
2. When that rejection arrived mid-stream, the stream ended with no content
block and OmniRoute SYNTHESIZED a 502 (stream.ts, code "empty_response").
markAccountUnavailable treats any status >= 500 on a per-model-quota
provider as an upstream per-model server error and locked claude-opus-5
for 60s — on a perfectly healthy account. With the other two accounts
genuinely rate-limited, all four targets were gone and the combo returned
the pre-dispatch 429. A 500 was already exempt; this 502 is not even an
upstream status, so exempt it too.
3. Every in-stream failure collapsed to "streaming upstream error", and the
client saw only "Claude returned an empty response (no content block)".
That masking hid both real rejections above. Carry the upstream error type
and message into the failure reason and the call log.
Measured 2026-09-21: claude-opus-5 was failing ~58% of native Claude Code
passthrough requests (63 of 109 on one key), against 79% success overall.
…-passthrough-empty-response
The native claude passthrough stripped the top-level `safeguards` field unconditionally to stop `400 safeguards: Extra inputs are not permitted`. That 400 only happens when the field reaches Anthropic without its paired `dangerous-tool-use-2026-09-03` beta; stripping it always made every gateway session ineligible for server-side auto mode (https://code.claude.com/docs/en/auto-mode-classifier-billing). Decide with the same mergeClientAnthropicBeta the executor uses: keep `safeguards` when the paired beta will be forwarded, strip it otherwise. The strip and the existing temperature/top_p strip move into stripClaudeRejectedTopLevelFields so chatCore.ts stays under its frozen size, and isSyntheticEmptyStreamFailure moves out of auth.ts for the same reason (also fixing its orphaned doc comment).
|
Update (d8b9885): the The original 400 ( What changed
Tests
Live evidence (operator production gateway, deployed ahead of merge together with #14694) One Linux release build was deployed to both nodes. They are byte-identical: per-file digest
|
b3867e4
into
diegosouzapw:release/v3.8.51
… reasons diegosouzapw#13630 gated the same-target retry after a pre-content streaming upstream error on an exact match: handlePreContentStreamRetry() (open-sse/services/combo/executeTargetClassify.ts:35) compared quality.reason !== "streaming upstream error". diegosouzapw#14314 (b3867e4) then intentionally appended the upstream detail to that reason in validateResponseQuality() ("streaming upstream error: <detail>") so the real Anthropic rejection reaches the call log. Every detailed reason failed the exact match, so protected (fallbackOnlyOnQuotaExhaustion) and native-pinned targets stopped retrying once and failed instead. validateQuality.ts now owns STREAMING_UPSTREAM_ERROR_REASON, builds the reason in one helper and exports isStreamingUpstreamErrorReason(), which the retry gate uses (bare or "<base>: <detail>", nothing else). The two reason assertions in the diegosouzapw#13630 test move to the exact new text diegosouzapw#14314 documents, and a classifier test pins both forms plus near-misses. Refs diegosouzapw#14547
An empty Claude passthrough stream no longer empties the whole combo. The chatCore.ts wiring for this change is carried by the agent sessions commit, which owns the final version of that shared file.
…egosouzapw#14864 agent sessions Agent session and project attribution (diegosouzapw#14833), the /v1/me/sessions self-service endpoints (diegosouzapw#14840), and /v1/me/sessions/{id}/messages with opt-in turn capture (diegosouzapw#14864), including the chatCore wiring, as deployed. Migrations keep the deployed numbers 9191-9193, which production databases have already applied. chatCore.ts also carries the diegosouzapw#14314 and diegosouzapw#14862 wiring.
Symptom
A client with four Claude accounts repeatedly got:
Two of those four accounts were healthy and not in cooldown. Clearing the cooldowns changed nothing — the cooldowns were never the cause.
What was actually happening
The "reset after 50m" is the rate-limited accounts' quota window, which is what made this look like a cooldown problem.
Measured on 2026-09-20/21:
claude-opus-5succeeded 78.9% of the time overall (521 attempts), but only 35.8% on the native Claude Code passthrough keys — 57.8% returned the empty-response 502 (63 of 109 on one key). It tracks the request shape, not the account.Three compounding defects
1.
safeguardsis forwarded and rejected. Anthropic rejects unknown top-level fields outright —400 safeguards: Extra inputs are not permitted. Newer Claude Code builds sendsafeguards; the pure-passthrough path forwards the body verbatim, so every such request 400s. Stripped, following the existingtop_pprecedent a few lines above. The field is client-side only and the request is rejected whole rather than degraded, so dropping it is strictly better than failing.2. A synthesized 502 caused a 60s per-model lockout on healthy accounts. When the rejection arrives mid-stream the stream ends with no content block, and OmniRoute synthesizes a 502 (
stream.ts→emitClaudeEmptyStreamErrorAndAbort, codeempty_response).markAccountUnavailabletreats anystatus >= 500on a per-model-quota provider as an upstream per-model server error and lockedclaude-opus-5for 60s — on an account that was fine. With the other two accounts genuinely rate-limited, all four targets vanished and the combo returned the pre-dispatch 429.A bare
500was already exempt as "intermittent and NOT model-specific". This 502 is not even an upstream status, so it is now exempt too.502/503/504from a real upstream keep the existing lockout path (#6216).3. The real error was masked. Every in-stream failure collapsed to
"streaming upstream error", and the client only ever saw"Claude returned an empty response (no content block)". That masking hid both rejections above for days. The upstream error type and message now flow into the failure reason and the call log.Tests
New
tests/unit/claude-passthrough-empty-response.test.ts(7 tests), each TDD-verified against the unfixed code:safeguardsis listed and stripped while real payload survivesvalidateQuality.tsreverted fails withreason must name the upstream error type, got: streaming upstream errorquality-validation-benign-error,combo-quality-validator-reasoning,combo-quality-tiny-budget-probe,anthropic-thinking-signature-recoverycombo-provider-cooldown-sibling(pins the 500 carve-out contract)combo-model-lockout-honors-reset-1308,combo-provider-cooldown,gemini-deprecated-model-lockouttypecheck:coreeslint(changed files)endpointPathwarning, present unpatched)prettier --checkThe
tool_additionhalf of this incident is a separate one-line fix on #13994, which owns the inline-tools code.