Skip to content

fix(combo): surface context-overflow before compression so oversized requests fail fast with a clear error (#10225) - #10503

Merged
diegosouzapw merged 2 commits into
release/v3.8.50from
fix/10225-combo-context-overflow
Aug 18, 2026
Merged

diegosouzapw merged 2 commits into
release/v3.8.50from
fix/10225-combo-context-overflow

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

Closes #10225

Combo requests whose combined context exceeds the model/context budget previously reached compression and yielded a late, generic failure instead of a clear context-overflow rejection.

Root cause: the context-overflow check ran after (or was bypassed by) the compression path, so an oversized request's failure was ambiguous and late.

Fix: check knownContextOverflow/context budget during dispatchPrelude and targetResolution, surfaced through combo.ts`` + chat.ts` as an early, explicit error before compression runs.

Regression test (RED→GREEN): tests/unit/combo-context-overflow-compression-probe.test.ts — 3 assertions GREEN (oversized request rejected explicitly pre-compression).

Gates: typecheck exit 0 · eslint exit 0 · file-size OK · complexity OK · cognitive OK · changelog-integrity OK · combo-context-overflow probe 3/3 GREEN.

⚠️ base-red inherited: #9985

@diegosouzapw

Copy link
Copy Markdown
Owner Author

Request changes. The new preflight deferral is not aligned with chatCore's actual compression eligibility: for a codex Responses target, shouldUseNativeCodexPassthrough() is true, and chatCore sets compressionExcluded from nativeCodexPassthrough, disabling both prompt compression and reactive compaction. However, resolveComboContextOverflowDeferral() returns defer=true based only on the global switch and API-key opt-out, while getKnownContextOverflow() checks only configured exclusions. A raw-over-limit codex request therefore skips the old early 400, reaches chatCore without compression, and is rejected by the final raw context gate rather than being compacted. The added probe passes because handleSingleModel is mocked to return 200 and never enters chatCore. Focused evidence: the PR probe passes 3/3; a head probe prints {"native":true,"deferralCondition":true}. Please make deferral target-aware and add a real chatCore regression for both compressed-success and compressed-still-too-large cases. The base-red #9985 remains inherited and is not attributed to this PR.

…dex passthrough (#10225)

The deferral added by the prior commit checked only operator-named
compression exclusions when deciding whether at least one target "can
compress" — it never accounted for native Codex Responses passthrough
targets, which chatCore.ts unconditionally excludes from compression
(compressionExcluded = nativeCodexPassthrough || ...). Deferring on such
a target's account let an oversized request skip both the combo preflight
AND compression, reaching fetch() uncompressed.

Thread the same request-shape facts chatCore.ts uses
(shouldUseNativeCodexPassthrough: provider/sourceFormat/endpointPath/body/
headers) down into getKnownContextOverflow so the deferral decision can
never drift from chatCore's own — a native-codex-passthrough target now
never counts as "compressible", so a pool made only of such targets keeps
the fast local 400 instead of a wasted round trip.

Adds regression coverage: the pure getKnownContextOverflow target-aware
check, an end-to-end handleComboChat proof that a native-codex-only pool
fails fast with zero dispatches, and two real handleChatCore-path tests
proving compression actually reduces the dispatched body when eligible,
and that a still-too-large-after-compression request is rejected locally
without an upstream call.
@diegosouzapw
diegosouzapw merged commit 9500adb into release/v3.8.50 Aug 18, 2026
11 checks passed
diegosouzapw added a commit to branben/OmniRoute that referenced this pull request Aug 18, 2026
Resync against 87 new release commits. Conflicts were all additive
option-bag fields landed in parallel by unrelated PRs — union both
sides in each case (no functional overlap):

- open-sse/services/combo.ts, combo/dispatchPrelude.ts, combo/types.ts:
  this PR's perTargetAdmission (diegosouzapw#9654 lane-aware admission probe)
  alongside release's deferContextOverflowWhenCompressible/
  compressionExclusions/sourceFormat/endpointPath/requestHeaders
  (diegosouzapw#10225/diegosouzapw#10503 compression deferral) — both sets of fields now
  coexist on HandleComboChatOptions and thread through unchanged.
- config/quality/open-sse-typecheck-baseline.json: unioned both
  sides' newly-frozen pre-existing typecheck entries (no overlapping
  file keys).
- tests/unit/feature-flags-settings.test.ts: the marked conflict was
  cosmetic wording only (hardcoded "49" vs the already-declared
  EXPECTED_FEATURE_FLAG_COUNT template), but the merge also silently
  combined two non-adjacent, non-conflicting array insertions in
  featureFlagDefinitions.ts — release's own flag growth (48→49) plus
  this PR's OMNIROUTE_CHAT_VIRTUAL_LANES flag — for a true post-merge
  total of 50. Bumped EXPECTED_FEATURE_FLAG_COUNT 49 -> 50 to match.

Verified: typecheck:core clean; the three diegosouzapw#9654 focused suites
(combo-lane-awareness, per-connection-admission, admission-virtual-lanes)
plus feature-flags-settings.test.ts all pass (81/81).

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
@diegosouzapw
diegosouzapw deleted the fix/10225-combo-context-overflow branch August 19, 2026 18:47
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 20, 2026
…requests fail fast with a clear error (diegosouzapw#10225) (diegosouzapw#10503)

* fix(combo): surface context-overflow before compression so oversized requests fail fast with a clear error (diegosouzapw#10225)

* fix(combo): make context-overflow deferral target-aware for native Codex passthrough (diegosouzapw#10225)

The deferral added by the prior commit checked only operator-named
compression exclusions when deciding whether at least one target "can
compress" — it never accounted for native Codex Responses passthrough
targets, which chatCore.ts unconditionally excludes from compression
(compressionExcluded = nativeCodexPassthrough || ...). Deferring on such
a target's account let an oversized request skip both the combo preflight
AND compression, reaching fetch() uncompressed.

Thread the same request-shape facts chatCore.ts uses
(shouldUseNativeCodexPassthrough: provider/sourceFormat/endpointPath/body/
headers) down into getKnownContextOverflow so the deferral decision can
never drift from chatCore's own — a native-codex-passthrough target now
never counts as "compressible", so a pool made only of such targets keeps
the fast local 400 instead of a wasted round trip.

Adds regression coverage: the pure getKnownContextOverflow target-aware
check, an end-to-end handleComboChat proof that a native-codex-only pool
fails fast with zero dispatches, and two real handleChatCore-path tests
proving compression actually reduces the dispatched body when eligible,
and that a still-too-large-after-compression request is rejected locally
without an upstream call.

---------

Co-authored-by: adevwithpurpose <adevwithpurpose@users.noreply.github.com>
giauphan pushed a commit to giauphan/OmniRoute that referenced this pull request Aug 20, 2026
…requests fail fast with a clear error (diegosouzapw#10225) (diegosouzapw#10503)

* fix(combo): surface context-overflow before compression so oversized requests fail fast with a clear error (diegosouzapw#10225)

* fix(combo): make context-overflow deferral target-aware for native Codex passthrough (diegosouzapw#10225)

The deferral added by the prior commit checked only operator-named
compression exclusions when deciding whether at least one target "can
compress" — it never accounted for native Codex Responses passthrough
targets, which chatCore.ts unconditionally excludes from compression
(compressionExcluded = nativeCodexPassthrough || ...). Deferring on such
a target's account let an oversized request skip both the combo preflight
AND compression, reaching fetch() uncompressed.

Thread the same request-shape facts chatCore.ts uses
(shouldUseNativeCodexPassthrough: provider/sourceFormat/endpointPath/body/
headers) down into getKnownContextOverflow so the deferral decision can
never drift from chatCore's own — a native-codex-passthrough target now
never counts as "compressible", so a pool made only of such targets keeps
the fast local 400 instead of a wasted round trip.

Adds regression coverage: the pure getKnownContextOverflow target-aware
check, an end-to-end handleComboChat proof that a native-codex-only pool
fails fast with zero dispatches, and two real handleChatCore-path tests
proving compression actually reduces the dispatched body when eligible,
and that a still-too-large-after-compression request is rejected locally
without an upstream call.

---------

Co-authored-by: adevwithpurpose <adevwithpurpose@users.noreply.github.com>
backryun added a commit to backryun/OmniRoute that referenced this pull request Aug 21, 2026
…iegosouzapw#10162

PR diegosouzapw#10162 (RFC diegosouzapw#10141 owner decision) made approximate gateway token
estimates advisory-only and deleted the pre-dispatch hard gate —
open-sse/services/combo/knownContextOverflow.ts and its
getKnownContextOverflow() helper no longer exist. Five tests in
combo-context-overflow-compression-probe.test.ts still imported and
called the removed helper, so every unit shard failed with
'TypeError: getKnownContextOverflow is not a function' plus a stale
'combo keeps the fast 400' case asserting the pre-diegosouzapw#10162 behavior
(got 502 after the gate removal).

Removed the five dead tests; kept the three that still pin live
behavior and pass:
- diegosouzapw#10225 combo does not early-400 a compressible request (deferral flag)
- diegosouzapw#10503 real chatCore path dispatches the COMPRESSED body upstream
- diegosouzapw#10503 still-too-large-after-compression rejects locally, zero dispatch

Added a file-header note explaining the diegosouzapw#10225/diegosouzapw#10503 -> diegosouzapw#10162 history
so the next reader knows why the pure-helper coverage is gone. Verified:
3/3 green locally.
backryun added a commit to backryun/OmniRoute that referenced this pull request Aug 21, 2026
…iegosouzapw#10162

PR diegosouzapw#10162 (RFC diegosouzapw#10141 owner decision) made approximate gateway token
estimates advisory-only and deleted the pre-dispatch hard gate —
open-sse/services/combo/knownContextOverflow.ts and its
getKnownContextOverflow() helper no longer exist. Five tests in
combo-context-overflow-compression-probe.test.ts still imported and
called the removed helper, so every unit shard failed with
'TypeError: getKnownContextOverflow is not a function' plus a stale
'combo keeps the fast 400' case asserting the pre-diegosouzapw#10162 behavior
(got 502 after the gate removal).

Removed the five dead tests; kept the three that still pin live
behavior and pass:
- diegosouzapw#10225 combo does not early-400 a compressible request (deferral flag)
- diegosouzapw#10503 real chatCore path dispatches the COMPRESSED body upstream
- diegosouzapw#10503 still-too-large-after-compression rejects locally, zero dispatch

Added a file-header note explaining the diegosouzapw#10225/diegosouzapw#10503 -> diegosouzapw#10162 history
so the next reader knows why the pure-helper coverage is gone. Verified:
3/3 green locally.
diegosouzapw added a commit that referenced this pull request Sep 24, 2026
…s lossy policy; drain 12 base-red files

#14529 moved lossy engines off the header-less path in the runtime but not in
deriveEffectivePreviewPlan(), so the dashboard preview showed rtk -> caveman
while requests ran session-dedup -> lite (#12063 again). The preview now
applies downgradeUnrequestedLossy(); the #12063 test pins preview == runtime
for the profile, engines-map and safe-profile cases (fails on the old code).

Test guards realigned to merged design changes, each traced to its commit:
- #14529: compression suites opt in with allow-lossy where they test the
  operator plan, and assert the header-less downgrade; default seed is
  session-dedup + lite; the #10503 probe ignores the loopback dashboard
  telemetry POST that stacked compression emits (the context gate still
  rejects locally — verified with logging).
- #14530: breaker profile asserted on the connection breaker; the
  provider-wide cooldown case drives the network-error path.
- #14370: @opencode/plugin (stable OpenCode 2.x contract, npm maintainer
  thdxr) added to the dependency allowlist.
- #14213: the opt-in empty-turn retry credential site is inventoried as
  class A (dispatch goes through assertManagedLeaseFence).
- #14223 / #14329: flag count 77, quota registry 13 (muse-code).

Refs #14496
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…requests fail fast with a clear error (diegosouzapw#10225) (diegosouzapw#10503)

* fix(combo): surface context-overflow before compression so oversized requests fail fast with a clear error (diegosouzapw#10225)

* fix(combo): make context-overflow deferral target-aware for native Codex passthrough (diegosouzapw#10225)

The deferral added by the prior commit checked only operator-named
compression exclusions when deciding whether at least one target "can
compress" — it never accounted for native Codex Responses passthrough
targets, which chatCore.ts unconditionally excludes from compression
(compressionExcluded = nativeCodexPassthrough || ...). Deferring on such
a target's account let an oversized request skip both the combo preflight
AND compression, reaching fetch() uncompressed.

Thread the same request-shape facts chatCore.ts uses
(shouldUseNativeCodexPassthrough: provider/sourceFormat/endpointPath/body/
headers) down into getKnownContextOverflow so the deferral decision can
never drift from chatCore's own — a native-codex-passthrough target now
never counts as "compressible", so a pool made only of such targets keeps
the fast local 400 instead of a wasted round trip.

Adds regression coverage: the pure getKnownContextOverflow target-aware
check, an end-to-end handleComboChat proof that a native-codex-only pool
fails fast with zero dispatches, and two real handleChatCore-path tests
proving compression actually reduces the dispatched body when eligible,
and that a still-too-large-after-compression request is rejected locally
without an upstream call.

---------

Co-authored-by: adevwithpurpose <adevwithpurpose@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(combo): defer known context overflow rejection until compression for Responses requests

2 participants