Repository navigation
fix(combo): surface context-overflow before compression so oversized requests fail fast with a clear error (#10225) - #10503
Conversation
…requests fail fast with a clear error (#10225)
|
Request changes. The new preflight deferral is not aligned with chatCore's actual compression eligibility: for a codex Responses target, shouldUseNativeCodexPassthrough() is true, and chatCore sets compressionExcluded from nativeCodexPassthrough, disabling both prompt compression and reactive compaction. However, resolveComboContextOverflowDeferral() returns defer=true based only on the global switch and API-key opt-out, while getKnownContextOverflow() checks only configured exclusions. A raw-over-limit codex request therefore skips the old early 400, reaches chatCore without compression, and is rejected by the final raw context gate rather than being compacted. The added probe passes because handleSingleModel is mocked to return 200 and never enters chatCore. Focused evidence: the PR probe passes 3/3; a head probe prints {"native":true,"deferralCondition":true}. Please make deferral target-aware and add a real chatCore regression for both compressed-success and compressed-still-too-large cases. The base-red #9985 remains inherited and is not attributed to this PR. |
…dex passthrough (#10225) The deferral added by the prior commit checked only operator-named compression exclusions when deciding whether at least one target "can compress" — it never accounted for native Codex Responses passthrough targets, which chatCore.ts unconditionally excludes from compression (compressionExcluded = nativeCodexPassthrough || ...). Deferring on such a target's account let an oversized request skip both the combo preflight AND compression, reaching fetch() uncompressed. Thread the same request-shape facts chatCore.ts uses (shouldUseNativeCodexPassthrough: provider/sourceFormat/endpointPath/body/ headers) down into getKnownContextOverflow so the deferral decision can never drift from chatCore's own — a native-codex-passthrough target now never counts as "compressible", so a pool made only of such targets keeps the fast local 400 instead of a wasted round trip. Adds regression coverage: the pure getKnownContextOverflow target-aware check, an end-to-end handleComboChat proof that a native-codex-only pool fails fast with zero dispatches, and two real handleChatCore-path tests proving compression actually reduces the dispatched body when eligible, and that a still-too-large-after-compression request is rejected locally without an upstream call.
Resync against 87 new release commits. Conflicts were all additive option-bag fields landed in parallel by unrelated PRs — union both sides in each case (no functional overlap): - open-sse/services/combo.ts, combo/dispatchPrelude.ts, combo/types.ts: this PR's perTargetAdmission (diegosouzapw#9654 lane-aware admission probe) alongside release's deferContextOverflowWhenCompressible/ compressionExclusions/sourceFormat/endpointPath/requestHeaders (diegosouzapw#10225/diegosouzapw#10503 compression deferral) — both sets of fields now coexist on HandleComboChatOptions and thread through unchanged. - config/quality/open-sse-typecheck-baseline.json: unioned both sides' newly-frozen pre-existing typecheck entries (no overlapping file keys). - tests/unit/feature-flags-settings.test.ts: the marked conflict was cosmetic wording only (hardcoded "49" vs the already-declared EXPECTED_FEATURE_FLAG_COUNT template), but the merge also silently combined two non-adjacent, non-conflicting array insertions in featureFlagDefinitions.ts — release's own flag growth (48→49) plus this PR's OMNIROUTE_CHAT_VIRTUAL_LANES flag — for a true post-merge total of 50. Bumped EXPECTED_FEATURE_FLAG_COUNT 49 -> 50 to match. Verified: typecheck:core clean; the three diegosouzapw#9654 focused suites (combo-lane-awareness, per-connection-admission, admission-virtual-lanes) plus feature-flags-settings.test.ts all pass (81/81). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…requests fail fast with a clear error (diegosouzapw#10225) (diegosouzapw#10503) * fix(combo): surface context-overflow before compression so oversized requests fail fast with a clear error (diegosouzapw#10225) * fix(combo): make context-overflow deferral target-aware for native Codex passthrough (diegosouzapw#10225) The deferral added by the prior commit checked only operator-named compression exclusions when deciding whether at least one target "can compress" — it never accounted for native Codex Responses passthrough targets, which chatCore.ts unconditionally excludes from compression (compressionExcluded = nativeCodexPassthrough || ...). Deferring on such a target's account let an oversized request skip both the combo preflight AND compression, reaching fetch() uncompressed. Thread the same request-shape facts chatCore.ts uses (shouldUseNativeCodexPassthrough: provider/sourceFormat/endpointPath/body/ headers) down into getKnownContextOverflow so the deferral decision can never drift from chatCore's own — a native-codex-passthrough target now never counts as "compressible", so a pool made only of such targets keeps the fast local 400 instead of a wasted round trip. Adds regression coverage: the pure getKnownContextOverflow target-aware check, an end-to-end handleComboChat proof that a native-codex-only pool fails fast with zero dispatches, and two real handleChatCore-path tests proving compression actually reduces the dispatched body when eligible, and that a still-too-large-after-compression request is rejected locally without an upstream call. --------- Co-authored-by: adevwithpurpose <adevwithpurpose@users.noreply.github.com>
…requests fail fast with a clear error (diegosouzapw#10225) (diegosouzapw#10503) * fix(combo): surface context-overflow before compression so oversized requests fail fast with a clear error (diegosouzapw#10225) * fix(combo): make context-overflow deferral target-aware for native Codex passthrough (diegosouzapw#10225) The deferral added by the prior commit checked only operator-named compression exclusions when deciding whether at least one target "can compress" — it never accounted for native Codex Responses passthrough targets, which chatCore.ts unconditionally excludes from compression (compressionExcluded = nativeCodexPassthrough || ...). Deferring on such a target's account let an oversized request skip both the combo preflight AND compression, reaching fetch() uncompressed. Thread the same request-shape facts chatCore.ts uses (shouldUseNativeCodexPassthrough: provider/sourceFormat/endpointPath/body/ headers) down into getKnownContextOverflow so the deferral decision can never drift from chatCore's own — a native-codex-passthrough target now never counts as "compressible", so a pool made only of such targets keeps the fast local 400 instead of a wasted round trip. Adds regression coverage: the pure getKnownContextOverflow target-aware check, an end-to-end handleComboChat proof that a native-codex-only pool fails fast with zero dispatches, and two real handleChatCore-path tests proving compression actually reduces the dispatched body when eligible, and that a still-too-large-after-compression request is rejected locally without an upstream call. --------- Co-authored-by: adevwithpurpose <adevwithpurpose@users.noreply.github.com>
…iegosouzapw#10162 PR diegosouzapw#10162 (RFC diegosouzapw#10141 owner decision) made approximate gateway token estimates advisory-only and deleted the pre-dispatch hard gate — open-sse/services/combo/knownContextOverflow.ts and its getKnownContextOverflow() helper no longer exist. Five tests in combo-context-overflow-compression-probe.test.ts still imported and called the removed helper, so every unit shard failed with 'TypeError: getKnownContextOverflow is not a function' plus a stale 'combo keeps the fast 400' case asserting the pre-diegosouzapw#10162 behavior (got 502 after the gate removal). Removed the five dead tests; kept the three that still pin live behavior and pass: - diegosouzapw#10225 combo does not early-400 a compressible request (deferral flag) - diegosouzapw#10503 real chatCore path dispatches the COMPRESSED body upstream - diegosouzapw#10503 still-too-large-after-compression rejects locally, zero dispatch Added a file-header note explaining the diegosouzapw#10225/diegosouzapw#10503 -> diegosouzapw#10162 history so the next reader knows why the pure-helper coverage is gone. Verified: 3/3 green locally.
…iegosouzapw#10162 PR diegosouzapw#10162 (RFC diegosouzapw#10141 owner decision) made approximate gateway token estimates advisory-only and deleted the pre-dispatch hard gate — open-sse/services/combo/knownContextOverflow.ts and its getKnownContextOverflow() helper no longer exist. Five tests in combo-context-overflow-compression-probe.test.ts still imported and called the removed helper, so every unit shard failed with 'TypeError: getKnownContextOverflow is not a function' plus a stale 'combo keeps the fast 400' case asserting the pre-diegosouzapw#10162 behavior (got 502 after the gate removal). Removed the five dead tests; kept the three that still pin live behavior and pass: - diegosouzapw#10225 combo does not early-400 a compressible request (deferral flag) - diegosouzapw#10503 real chatCore path dispatches the COMPRESSED body upstream - diegosouzapw#10503 still-too-large-after-compression rejects locally, zero dispatch Added a file-header note explaining the diegosouzapw#10225/diegosouzapw#10503 -> diegosouzapw#10162 history so the next reader knows why the pure-helper coverage is gone. Verified: 3/3 green locally.
…s lossy policy; drain 12 base-red files #14529 moved lossy engines off the header-less path in the runtime but not in deriveEffectivePreviewPlan(), so the dashboard preview showed rtk -> caveman while requests ran session-dedup -> lite (#12063 again). The preview now applies downgradeUnrequestedLossy(); the #12063 test pins preview == runtime for the profile, engines-map and safe-profile cases (fails on the old code). Test guards realigned to merged design changes, each traced to its commit: - #14529: compression suites opt in with allow-lossy where they test the operator plan, and assert the header-less downgrade; default seed is session-dedup + lite; the #10503 probe ignores the loopback dashboard telemetry POST that stacked compression emits (the context gate still rejects locally — verified with logging). - #14530: breaker profile asserted on the connection breaker; the provider-wide cooldown case drives the network-error path. - #14370: @opencode/plugin (stable OpenCode 2.x contract, npm maintainer thdxr) added to the dependency allowlist. - #14213: the opt-in empty-turn retry credential site is inventoried as class A (dispatch goes through assertManagedLeaseFence). - #14223 / #14329: flag count 77, quota registry 13 (muse-code). Refs #14496
…requests fail fast with a clear error (diegosouzapw#10225) (diegosouzapw#10503) * fix(combo): surface context-overflow before compression so oversized requests fail fast with a clear error (diegosouzapw#10225) * fix(combo): make context-overflow deferral target-aware for native Codex passthrough (diegosouzapw#10225) The deferral added by the prior commit checked only operator-named compression exclusions when deciding whether at least one target "can compress" — it never accounted for native Codex Responses passthrough targets, which chatCore.ts unconditionally excludes from compression (compressionExcluded = nativeCodexPassthrough || ...). Deferring on such a target's account let an oversized request skip both the combo preflight AND compression, reaching fetch() uncompressed. Thread the same request-shape facts chatCore.ts uses (shouldUseNativeCodexPassthrough: provider/sourceFormat/endpointPath/body/ headers) down into getKnownContextOverflow so the deferral decision can never drift from chatCore's own — a native-codex-passthrough target now never counts as "compressible", so a pool made only of such targets keeps the fast local 400 instead of a wasted round trip. Adds regression coverage: the pure getKnownContextOverflow target-aware check, an end-to-end handleComboChat proof that a native-codex-only pool fails fast with zero dispatches, and two real handleChatCore-path tests proving compression actually reduces the dispatched body when eligible, and that a still-too-large-after-compression request is rejected locally without an upstream call. --------- Co-authored-by: adevwithpurpose <adevwithpurpose@users.noreply.github.com>
Closes #10225
Combo requests whose combined context exceeds the model/context budget previously reached compression and yielded a late, generic failure instead of a clear context-overflow rejection.
Root cause: the context-overflow check ran after (or was bypassed by) the compression path, so an oversized request's failure was ambiguous and late.
Fix: check
knownContextOverflow/context budget duringdispatchPreludeandtargetResolution, surfaced throughcombo.ts`` +chat.ts` as an early, explicit error before compression runs.Regression test (RED→GREEN):
tests/unit/combo-context-overflow-compression-probe.test.ts— 3 assertions GREEN (oversized request rejected explicitly pre-compression).Gates: typecheck exit 0 · eslint exit 0 · file-size OK · complexity OK · cognitive OK · changelog-integrity OK · combo-context-overflow probe 3/3 GREEN.