Skip to content

fix(routing): bare model ids route to codex first; validate synced candidates - #9275

Merged
diegosouzapw merged 2 commits into
release/v3.8.50from
fix/bare-model-routing-precedence
Aug 3, 2026
Merged

diegosouzapw merged 2 commits into
release/v3.8.50from
fix/bare-model-routing-precedence

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

Two bare-model-routing bugs surfaced in the field when an OmniRoute deployment had a codex subscription whose cookie quota was exhausted (retry-after 429047s / ~5 days) AND an active kiro connection whose upstream sync briefly advertised 'claude-opus-5' before kiro vendored it into the static registry.

Bug 1: Bare 'gpt-5.6-sol' routes to codex even when codex is exhausted

Bare 'gpt-5.6-sol' (and friends) routed to the codex provider via the 'codex-only installations' branch of resolveModelByProviderInference — even when the user had explicitly configured 'agentrouter' as their model_provider in the codex CLI config. With codex in cooldown (retry-after: 429047s ≈ 5 days), every bare request 429'd.

Fix: extend CODEX_NATIVE_UNPREFIXED_MODELS to include the full gpt-5.6-sol tier set + gpt-5.5 + the related codex-native ids. The Codex CLI default is now actually honored; users can still prefix 'agentrouter/gpt-5.6-sol' to opt into a specific provider (existing regression-tested).

Bug 2: Bare 'claude-opus-5' silently routes to kiro (synced-catalog phantom)

Bare 'claude-opus-5' silently routed to 'kiro' when kiro's synced /v1/models catalog had that id (transient upstream quirk). kiro's static registry never cataloged claude-opus-5, so the upstream call 404'd with 'No active credentials for provider: kiro'.

Fix: validate activeSyncedProviders against MODEL_TO_PROVIDERS before merging them into the candidate list. Auto-discovery still wins when the model id has no static entry (brand-new models from upstream keep working).

Bonus: error message now includes candidate-prefix hint

When handleNoCredentials returns a 404 'No active credentials for provider: X', surface the top-3 candidate aliases (e.g. 'anthropic/claude-opus-5, claude/claude-opus-5, agentrouter/claude-opus-5') so the operator can pick a working prefix instead of staring at a wall.

Files changed

  • open-sse/services/model.ts — CODEX_NATIVE_UNPREFIXED_MODELS expanded + synced validation
  • src/sse/handlers/chatHelpers.ts — handleNoCredentials accepts candidateAliases
  • src/sse/handlers/chat.ts — passes resolved.candidateAliases to handleNoCredentials
  • tests/unit/fix-bare-model-precedence.test.ts (7 tests)
  • tests/unit/fix-synced-model-validation.test.ts (3 tests)
  • tests/unit/fix-error-message-candidates.test.ts (3 tests)
  • tests/unit/fix-bare-routing-fallback.test.ts (7 tests)

Validation

  • npm run typecheck:core → ✅
  • All 4 new test files: 20 tests, all pass
  • model-resolver.test.ts: 25 tests, all pass (regression)
  • bare-model-autopick.test.ts: 13 tests, all pass (regression)
  • agentrouter-chatcore-protocols.test.ts + agentrouter-models-discovery-7016.test.ts: all pass
  • chat-cooldown-aware-retry.test.ts: all pass
  • npm run lint: 2 pre-existing errors in @/lib/localDb imports (untouched by this PR — already in eslint-suppressions.json)

…ndidates

Two bare-model-routing bugs surfaced in the field when an OmniRoute
deployment had a codex subscription whose cookie quota was exhausted
(retry-after 429047s / ~5 days) AND an active kiro connection whose
upstream sync briefly advertised 'claude-opus-5' before kiro vendored
it into the static registry.

  1. Bare 'gpt-5.6-sol' (and friends) routed to the codex provider even
     when the user had explicitly configured 'agentrouter' as their
     provider (via model_provider in codex CLI). With codex in cooldown,
     every bare request 429'd. Fix: extend CODEX_NATIVE_UNPREFIXED_MODELS
     to include the full gpt-5.6-sol tier set + gpt-5.5 + the related
     codex-native ids. The Codex CLI default is now actually honored;
     users can still prefix 'agentrouter/gpt-5.6-sol' to opt into a
     specific provider.

  2. Bare 'claude-opus-5' silently routed to 'kiro' when kiro's synced
     /v1/models catalog had that id (likely from a transient upstream
     quirk). kiro's static registry never cataloged claude-opus-5, so
     the upstream call 404'd. Fix: validate activeSyncedProviders against
     MODEL_TO_PROVIDERS before merging them into the candidate list.
     Auto-discovery still wins when the model id has no static entry
     (brand-new models from upstream keep working).

Bonus: when handleNoCredentials returns a 404 'No active credentials for
provider: X' error, surface the top-3 candidate aliases (e.g.
'anthropic/claude-opus-5, claude/claude-opus-5, agentrouter/claude-opus-5')
so the operator can pick a working prefix instead of staring at a wall.

Tests (all pass, 25 regression tests preserved):
  - tests/unit/fix-bare-model-precedence.test.ts (7 tests)
  - tests/unit/fix-synced-model-validation.test.ts (3 tests)
  - tests/unit/fix-error-message-candidates.test.ts (3 tests)
  - tests/unit/fix-bare-routing-fallback.test.ts (7 tests)
…r WAF

The agentrouter.org WAF blocks requests containing 'lorem ipsum' in
messages[].content. When Claude Code reads test files via the Read tool,
the content appears in tool_result blocks which can trigger the filter.

Replace 'lorem ipsum dolor sit amet' with 'example content for testing
purposes' in compression harness test to avoid false positives.
@diegosouzapw
diegosouzapw merged commit a72e165 into release/v3.8.50 Aug 3, 2026
10 checks passed
diegosouzapw added a commit that referenced this pull request Aug 3, 2026
The agentrouter.org upstream WAF returns 400 content-blocked
intermittently when:
  1. messages[].content contains a blocked keyword (Lorem ipsum, the
     phrase 'language model' alone, 'virtual assistant', etc.); or
  2. Requests from the same IP/key arrive in a burst, after which the
     WAF's per-IP suspicion bucket starts blocking content that would
     normally pass. The bucket relaxes after ~5-10s of idle.

Apply three mitigations:

1. Burst guard (open-sse/services/wafRateLimit.ts)
   Per-bucket (provider+url) gate that enforces a 500ms minimum gap
   between outbound requests to agentrouter. Configurable via
   configureWafRateLimit(). Tested in tests/unit/wafRateLimit.test.ts.

2. Reactive retry (BaseExecutor.WAF_RETRY_CONFIG in base.ts)
   New WAF_RETRY_CONFIG with maxAttempts=2, delayMs=1500,
   backoffMultiplier=2. When the upstream returns 400 with a body that
   matches /content[_-]blocked/i, retry the same URL with exponential
   backoff (1.5s, 3.0s) before falling through to the 429/401/fallback
   chain. Tested in tests/unit/base-executor-waf-retry.test.ts.

3. Documentation (docs/security/AGENTROUTER_WAF.md)
   Blocklist of always-blocked and almost-always-blocked patterns,
   behavior under load, guidance for prompts/tool output, and pointers
   to the relevant code paths in OmniRoute.

These are belt-and-suspenders: the burst guard prevents the WAF from
activating on normal traffic, and the reactive retry recovers when it
does anyway. Together they should eliminate the intermittent
400 content-blocked that Claude Code sees when running through
agentrouter via OmniRoute.

Refs #9275 follow-up. Test: 'WAF retry config shape' and 'WAF retry
differs from generic' guard the WAF_RETRY_CONFIG contract so future
refactors don't accidentally collapse the two retry paths.

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
wgordon17 added a commit to wgordon17/OmniRoute that referenced this pull request Aug 4, 2026
Two categories of inherited base-branch breakage surfaced when
rebasing onto release/v3.8.50's latest tip, both confirmed unrelated
to this PR's own diff:

- check:file-size: base.ts and chat.ts drifted further past their
  frozen caps via already-merged commits (7163081 and others) that
  didn't rebaseline after growing them. Documented and bumped in
  file-size-baseline.json.
- chat-helpers.test.ts: two gpt-5.5 routing assertions predate diegosouzapw#9275
  (fix(routing): bare model ids route to codex first), which
  deliberately made gpt-5.5 route to codex unconditionally, regardless
  of which other providers are active. Confirmed via diegosouzapw#9275's own
  commit message and code comments this is intentional, not a
  regression; verified reproducible on the raw base tip alone, with
  no changes from this PR involved. Updated both assertions and their
  names to match the new, intentional default.
wgordon17 added a commit to wgordon17/OmniRoute that referenced this pull request Aug 4, 2026
Same two categories of inherited base-branch breakage as diegosouzapw#9006 (this
PR shares the identical chatHelpers.ts diff), confirmed unrelated to
this PR's own diff:

- check:file-size: base.ts and chat.ts drifted further past their
  frozen caps via already-merged commits (7163081 and others) that
  didn't rebaseline after growing them.
- chat-helpers.test.ts: two gpt-5.5 routing assertions predate diegosouzapw#9275
  (fix(routing): bare model ids route to codex first), which
  deliberately made gpt-5.5 route to codex unconditionally. Updated
  both assertions and their names to match the new, intentional
  default.
wgordon17 added a commit to wgordon17/OmniRoute that referenced this pull request Aug 4, 2026
Previous commit fixed the regression by simply omitting gpt-5.5 from
CODEX_NATIVE_UNPREFIXED_MODELS, relying on fallthrough to the
pre-existing precedence logic below. That's structurally the same
failure mode that caused the original bug: a check "just happens" not
to apply to a given model id, with correctness depending on nobody
re-adding it later without reading a comment 20 lines away from the
call site (exactly what happened once already, in diegosouzapw#9275).

Adds a second, explicitly-named set,
CODEX_NATIVE_MODELS_WITH_OPENAI_PRECEDENCE, and checks it directly at
the call site: `CODEX_NATIVE_UNPREFIXED_MODELS.has(modelId) &&
!CODEX_NATIVE_MODELS_WITH_OPENAI_PRECEDENCE.has(modelId)`. A future
engineer reading this exact line sees the carve-out and why it exists,
instead of needing to notice gpt-5.5's absence from a list defined
well above it. Test updated to assert membership in the new set
directly, rather than only asserting absence from the old one.
diegosouzapw added a commit that referenced this pull request Aug 4, 2026
… guards

The two files #9275 added assert that bare gpt-5.5 / gpt-5.6-sol reach codex, but
they ran against an empty database — so they also pinned 'codex wins with no codex
connection at all', which is the regression #9447 removes. That put them in direct
contradiction with plan3-p0 / chat-helpers / codex-gpt55-routing-5887, which assert
openai for the very same input: no implementation could satisfy both, which is why
the release could not go green.

Seeding an active codex connection keeps the contract these files were written to
guard (codex beats openai for a Codex-native bare id) while dropping the accidental
'even with no codex configured' half. Cases that need no connection are left as they
were: the tier-only ids and codex-auto-review have no alternative provider to preempt,
and the explicit-prefix overrides are unaffected.
diegosouzapw added a commit that referenced this pull request Aug 4, 2026
…codex is active (#9447)

* fix(routing): only let Codex-native bare ids preempt a provider when codex is active

#9275 widened CODEX_NATIVE_UNPREFIXED_MODELS from a single id to gpt-5.5 plus the
gpt-5.6-sol/terra/luna tiers, so bare Codex CLI ids would reach the ChatGPT
subscription instead of fanning out to whichever provider won the inference race.
The early return it added never consulted the active-provider set, which made the
codex-only guard 30 lines below unreachable for every id in the set:

  if (CODEX_NATIVE_UNPREFIXED_MODELS.has(modelId)) return { provider: "codex", ... }

An OpenAI-only install therefore had bare gpt-5.5 routed to codex and failed with
'no active credentials for provider: codex' on a model OpenAI serves, and an install
whose codex connection was merely inactive failed identically. This also silently
reverted #5887's compatibility boundary.

The preference now only PREEMPTS another provider when a codex connection is active.
Ids that no other provider catalogs (codex-auto-review) still resolve to codex with no
connection at all — there is nothing to preempt and 'no codex credentials' is the
honest error. With codex active the preference still beats OpenAI, which is the point
of #9275, and an explicit openai/ prefix overrides it either way.

Tests: the three assertions that encode the intended #9275 change now expect codex
(plus a new one pinning the explicit-prefix override); the rest were already correct
and pass again untouched. Adds a regression test for the OpenAI-only case.

* docs(changelog): correct fragment id to #9447

* test(routing): seed an active codex connection in the bare-precedence guards

The two files #9275 added assert that bare gpt-5.5 / gpt-5.6-sol reach codex, but
they ran against an empty database — so they also pinned 'codex wins with no codex
connection at all', which is the regression #9447 removes. That put them in direct
contradiction with plan3-p0 / chat-helpers / codex-gpt55-routing-5887, which assert
openai for the very same input: no implementation could satisfy both, which is why
the release could not go green.

Seeding an active codex connection keeps the contract these files were written to
guard (codex beats openai for a Codex-native bare id) while dropping the accidental
'even with no codex configured' half. Cases that need no connection are left as they
were: the tier-only ids and codex-auto-review have no alternative provider to preempt,
and the explicit-prefix overrides are unaffected.

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
diegosouzapw pushed a commit that referenced this pull request Aug 4, 2026
…9392)

#9275 started appending a candidate-alias hint to the zero-active-credentials
error and terminated the provider name with a period, so the two sentences read
as one message. The two vscode tokenized-route tests still assert the old
unterminated string and now fail on every pull request opened against this
branch.

The Quality Gates workflow only runs on pull_request to release/**, never on
push, so the branch itself never re-runs these shards and the drift stayed
invisible after the merge.

Assert what the handler actually produces. Keeping the comparison exact rather
than loosening it to a prefix match is deliberate -- the exact form is what
caught the drift.

Signed-off-by: Minxi Hou <houminxi@gmail.com>
@diegosouzapw
diegosouzapw deleted the fix/bare-model-routing-precedence branch August 4, 2026 20:50
diegosouzapw added a commit to wgordon17/OmniRoute that referenced this pull request Aug 5, 2026
…gpt-5.5 contract (owner decision)

Keeps this PR's surviving value (the diegosouzapw#9275 comment fixes + the documented
diegosouzapw#5887×diegosouzapw#9447 interplay at the preemption site + own-growth baseline entry) and
drops the behavioral carve-out: the base's diegosouzapw#9447 active-connection bound already
delivers the agreed contract (OpenAI while codex is inactive, active codex
preempts, openai/ prefix overrides), encoded in codex-gpt55-routing-5887.test.ts.
Inherited file-size bumps reverted (release captain domain; diegosouzapw#9355 base-relative
mode protects innocent PRs).
diegosouzapw pushed a commit to wgordon17/OmniRoute that referenced this pull request Aug 11, 2026
Two categories of inherited base-branch breakage surfaced when
rebasing onto release/v3.8.50's latest tip, both confirmed unrelated
to this PR's own diff:

- check:file-size: base.ts and chat.ts drifted further past their
  frozen caps via already-merged commits (7163081 and others) that
  didn't rebaseline after growing them. Documented and bumped in
  file-size-baseline.json.
- chat-helpers.test.ts: two gpt-5.5 routing assertions predate diegosouzapw#9275
  (fix(routing): bare model ids route to codex first), which
  deliberately made gpt-5.5 route to codex unconditionally, regardless
  of which other providers are active. Confirmed via diegosouzapw#9275's own
  commit message and code comments this is intentional, not a
  regression; verified reproducible on the raw base tip alone, with
  no changes from this PR involved. Updated both assertions and their
  names to match the new, intentional default.
diegosouzapw pushed a commit to wgordon17/OmniRoute that referenced this pull request Aug 11, 2026
Two categories of inherited base-branch breakage surfaced when
rebasing onto release/v3.8.50's latest tip, both confirmed unrelated
to this PR's own diff:

- check:file-size: base.ts and chat.ts drifted further past their
  frozen caps via already-merged commits (7163081 and others) that
  didn't rebaseline after growing them. Documented and bumped in
  file-size-baseline.json.
- chat-helpers.test.ts: two gpt-5.5 routing assertions predate diegosouzapw#9275
  (fix(routing): bare model ids route to codex first), which
  deliberately made gpt-5.5 route to codex unconditionally, regardless
  of which other providers are active. Confirmed via diegosouzapw#9275's own
  commit message and code comments this is intentional, not a
  regression; verified reproducible on the raw base tip alone, with
  no changes from this PR involved. Updated both assertions and their
  names to match the new, intentional default.
diegosouzapw pushed a commit that referenced this pull request Aug 11, 2026
…n every provider (#9006)

* fix(executors): route Claude-via-Vertex through native rawPredict with real streaming

Claude models on Vertex AI were being sent through the generic OpenAI-
compatible partner endpoint, which 404s/errors for Claude on at least
some projects. Route them through Vertex's native Anthropic Messages
API (publishers/anthropic/.../rawPredict) instead, stripping the
body-level model field rawPredict rejects and injecting the required
anthropic_version field.

rawPredict only ever returns a complete JSON body, never real SSE
framing, so streaming requests now get a genuine Anthropic-format SSE
stream synthesized from that JSON (message_start/content_block_*/
message_delta/message_stop), which the existing claude-to-openai
response translator already knows how to parse.

Also fixes two response-format resolution bugs that silently dropped
a custom model's DB-stored targetFormat override whenever the model
id also existed in the static provider registry (as claude-sonnet-4-6
and claude-opus-4-7 do under vertex): resolveModelOrError had its own
ad-hoc resolution that never consulted the override, and even once
fixed, executeChatWithBreaker discarded the correctly-resolved format
before handleChatCore's own resolution ran a second time.

* docs: add changelog fragment for #8909

* refactor(sse): extract shared Claude effort-model predicate

* fix(sse): strip Claude effort-suffix ids for any provider serving a real Claude model

* fix(sse): keep no-think and CC-discovery catalog variant roots unprefixed

* fix(dashboard): re-qualify no-think playground model ids correctly

* fix(sse): scope Vertex 404s to a per-model lockout via passthroughModels

* docs: add changelog fragment for the Claude catalog/dispatch fix

* fix(sse): align regex naming and changelog formatting

* fix(sse): clarify effort-variant strip comment and add cross-module drift guard

* fix(sse): disambiguate Vertex connection-wide vs per-model 403s

* docs: document Vertex 403 disambiguation in changelog fragment

* fix(sse): correlate reason and resource within the same ErrorInfo detail

* fix(sse): extract Vertex error classifier and rebaseline frozen file sizes

* test: register vertex-passthrough-model-lockout in stryker tap.testFiles

* fix(sse): reconciles rebase-onto-tip drift for 9006

Two categories of inherited base-branch breakage surfaced when
rebasing onto release/v3.8.50's latest tip, both confirmed unrelated
to this PR's own diff:

- check:file-size: base.ts and chat.ts drifted further past their
  frozen caps via already-merged commits (7163081 and others) that
  didn't rebaseline after growing them. Documented and bumped in
  file-size-baseline.json.
- chat-helpers.test.ts: two gpt-5.5 routing assertions predate #9275
  (fix(routing): bare model ids route to codex first), which
  deliberately made gpt-5.5 route to codex unconditionally, regardless
  of which other providers are active. Confirmed via #9275's own
  commit message and code comments this is intentional, not a
  regression; verified reproducible on the raw base tip alone, with
  no changes from this PR involved. Updated both assertions and their
  names to match the new, intentional default.

* ci: re-trigger checks after GitHub Actions incident (2026-08-07, resolved)

* ci: re-trigger checks (previous push event was dropped)

* fix(quality): rebaseline combo-routing-engine.test.ts own-comment growth

The ALL_ACCOUNTS_INACTIVE->ALL_TARGETS_SKIPPED fix (a32aed7) added explanatory comments (+7 lines), pushing the file past its frozen 3457 cap. CI's PR-mode check:file-size caught it; local check-file-size.mjs was not re-run after that specific commit.
QuangBlue pushed a commit to QuangBlue/OmniRoute that referenced this pull request Sep 24, 2026
CODEX_NATIVE_UNPREFIXED_MODELS lists the Codex model ids that should reach
the ChatGPT subscription when a client sends them without a provider prefix
(diegosouzapw#9275), bounded to installs with an active Codex connection (diegosouzapw#9447). GPT-6
Astra came to the Codex catalog in diegosouzapw#13026 with the same effort tiers as
gpt-5.6-sol/terra, but its ids were never added to the set. With both a
Codex and an OpenAI connection active:

- bare `gpt-5.6-sol` routed to Codex, bare `gpt-6-astra` routed to OpenAI;
- /v1/models listed bare `gpt-5.6-*` ids but no bare `gpt-6-astra*` id.

Add gpt-6-astra and its six effort variants. The diegosouzapw#9447 bound still applies:
OpenAI's catalog also lists `gpt-6-astra`, so without an active Codex
connection the bare id keeps resolving to OpenAI. The effort variants are
Codex-family ids (codex and codex-app-server); they already routed to Codex
when it was active, so for them only the /v1/models rows change, exactly as
for the gpt-5.6 variants.

Test: tests/unit/codex-gpt6-astra-bare-id.test.ts (3 of 4 cases fail on the
previous model.ts; the fourth pins the no-Codex fallback). Three comments in
models-catalog-low-noise-flag.test.ts counted the set's size ("26 of the 27")
and one cited its old line range; they now describe the one exception
without a count and point at the file only.
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…ndidates (diegosouzapw#9275)

* fix(routing): bare model ids route to codex first; validate synced candidates

Two bare-model-routing bugs surfaced in the field when an OmniRoute
deployment had a codex subscription whose cookie quota was exhausted
(retry-after 429047s / ~5 days) AND an active kiro connection whose
upstream sync briefly advertised 'claude-opus-5' before kiro vendored
it into the static registry.

  1. Bare 'gpt-5.6-sol' (and friends) routed to the codex provider even
     when the user had explicitly configured 'agentrouter' as their
     provider (via model_provider in codex CLI). With codex in cooldown,
     every bare request 429'd. Fix: extend CODEX_NATIVE_UNPREFIXED_MODELS
     to include the full gpt-5.6-sol tier set + gpt-5.5 + the related
     codex-native ids. The Codex CLI default is now actually honored;
     users can still prefix 'agentrouter/gpt-5.6-sol' to opt into a
     specific provider.

  2. Bare 'claude-opus-5' silently routed to 'kiro' when kiro's synced
     /v1/models catalog had that id (likely from a transient upstream
     quirk). kiro's static registry never cataloged claude-opus-5, so
     the upstream call 404'd. Fix: validate activeSyncedProviders against
     MODEL_TO_PROVIDERS before merging them into the candidate list.
     Auto-discovery still wins when the model id has no static entry
     (brand-new models from upstream keep working).

Bonus: when handleNoCredentials returns a 404 'No active credentials for
provider: X' error, surface the top-3 candidate aliases (e.g.
'anthropic/claude-opus-5, claude/claude-opus-5, agentrouter/claude-opus-5')
so the operator can pick a working prefix instead of staring at a wall.

Tests (all pass, 25 regression tests preserved):
  - tests/unit/fix-bare-model-precedence.test.ts (7 tests)
  - tests/unit/fix-synced-model-validation.test.ts (3 tests)
  - tests/unit/fix-error-message-candidates.test.ts (3 tests)
  - tests/unit/fix-bare-routing-fallback.test.ts (7 tests)

* fix(tests): replace lorem ipsum with neutral text to avoid agentrouter WAF

The agentrouter.org WAF blocks requests containing 'lorem ipsum' in
messages[].content. When Claude Code reads test files via the Read tool,
the content appears in tool_result blocks which can trigger the filter.

Replace 'lorem ipsum dolor sit amet' with 'example content for testing
purposes' in compression harness test to avoid false positives.

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…uzapw#9323)

The agentrouter.org upstream WAF returns 400 content-blocked
intermittently when:
  1. messages[].content contains a blocked keyword (Lorem ipsum, the
     phrase 'language model' alone, 'virtual assistant', etc.); or
  2. Requests from the same IP/key arrive in a burst, after which the
     WAF's per-IP suspicion bucket starts blocking content that would
     normally pass. The bucket relaxes after ~5-10s of idle.

Apply three mitigations:

1. Burst guard (open-sse/services/wafRateLimit.ts)
   Per-bucket (provider+url) gate that enforces a 500ms minimum gap
   between outbound requests to agentrouter. Configurable via
   configureWafRateLimit(). Tested in tests/unit/wafRateLimit.test.ts.

2. Reactive retry (BaseExecutor.WAF_RETRY_CONFIG in base.ts)
   New WAF_RETRY_CONFIG with maxAttempts=2, delayMs=1500,
   backoffMultiplier=2. When the upstream returns 400 with a body that
   matches /content[_-]blocked/i, retry the same URL with exponential
   backoff (1.5s, 3.0s) before falling through to the 429/401/fallback
   chain. Tested in tests/unit/base-executor-waf-retry.test.ts.

3. Documentation (docs/security/AGENTROUTER_WAF.md)
   Blocklist of always-blocked and almost-always-blocked patterns,
   behavior under load, guidance for prompts/tool output, and pointers
   to the relevant code paths in OmniRoute.

These are belt-and-suspenders: the burst guard prevents the WAF from
activating on normal traffic, and the reactive retry recovers when it
does anyway. Together they should eliminate the intermittent
400 content-blocked that Claude Code sees when running through
agentrouter via OmniRoute.

Refs diegosouzapw#9275 follow-up. Test: 'WAF retry config shape' and 'WAF retry
differs from generic' guard the WAF_RETRY_CONFIG contract so future
refactors don't accidentally collapse the two retry paths.

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…codex is active (diegosouzapw#9447)

* fix(routing): only let Codex-native bare ids preempt a provider when codex is active

diegosouzapw#9275 widened CODEX_NATIVE_UNPREFIXED_MODELS from a single id to gpt-5.5 plus the
gpt-5.6-sol/terra/luna tiers, so bare Codex CLI ids would reach the ChatGPT
subscription instead of fanning out to whichever provider won the inference race.
The early return it added never consulted the active-provider set, which made the
codex-only guard 30 lines below unreachable for every id in the set:

  if (CODEX_NATIVE_UNPREFIXED_MODELS.has(modelId)) return { provider: "codex", ... }

An OpenAI-only install therefore had bare gpt-5.5 routed to codex and failed with
'no active credentials for provider: codex' on a model OpenAI serves, and an install
whose codex connection was merely inactive failed identically. This also silently
reverted diegosouzapw#5887's compatibility boundary.

The preference now only PREEMPTS another provider when a codex connection is active.
Ids that no other provider catalogs (codex-auto-review) still resolve to codex with no
connection at all — there is nothing to preempt and 'no codex credentials' is the
honest error. With codex active the preference still beats OpenAI, which is the point
of diegosouzapw#9275, and an explicit openai/ prefix overrides it either way.

Tests: the three assertions that encode the intended diegosouzapw#9275 change now expect codex
(plus a new one pinning the explicit-prefix override); the rest were already correct
and pass again untouched. Adds a regression test for the OpenAI-only case.

* docs(changelog): correct fragment id to diegosouzapw#9447

* test(routing): seed an active codex connection in the bare-precedence guards

The two files diegosouzapw#9275 added assert that bare gpt-5.5 / gpt-5.6-sol reach codex, but
they ran against an empty database — so they also pinned 'codex wins with no codex
connection at all', which is the regression diegosouzapw#9447 removes. That put them in direct
contradiction with plan3-p0 / chat-helpers / codex-gpt55-routing-5887, which assert
openai for the very same input: no implementation could satisfy both, which is why
the release could not go green.

Seeding an active codex connection keeps the contract these files were written to
guard (codex beats openai for a Codex-native bare id) while dropping the accidental
'even with no codex configured' half. Cases that need no connection are left as they
were: the tier-only ids and codex-auto-review have no alternative provider to preempt,
and the explicit-prefix overrides are unaffected.

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…iegosouzapw#9392)

diegosouzapw#9275 started appending a candidate-alias hint to the zero-active-credentials
error and terminated the provider name with a period, so the two sentences read
as one message. The two vscode tokenized-route tests still assert the old
unterminated string and now fail on every pull request opened against this
branch.

The Quality Gates workflow only runs on pull_request to release/**, never on
push, so the branch itself never re-runs these shards and the drift stayed
invisible after the merge.

Assert what the handler actually produces. Keeping the comparison exact rather
than loosening it to a prefix match is deliberate -- the exact form is what
caught the drift.

Signed-off-by: Minxi Hou <houminxi@gmail.com>
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…n every provider (diegosouzapw#9006)

* fix(executors): route Claude-via-Vertex through native rawPredict with real streaming

Claude models on Vertex AI were being sent through the generic OpenAI-
compatible partner endpoint, which 404s/errors for Claude on at least
some projects. Route them through Vertex's native Anthropic Messages
API (publishers/anthropic/.../rawPredict) instead, stripping the
body-level model field rawPredict rejects and injecting the required
anthropic_version field.

rawPredict only ever returns a complete JSON body, never real SSE
framing, so streaming requests now get a genuine Anthropic-format SSE
stream synthesized from that JSON (message_start/content_block_*/
message_delta/message_stop), which the existing claude-to-openai
response translator already knows how to parse.

Also fixes two response-format resolution bugs that silently dropped
a custom model's DB-stored targetFormat override whenever the model
id also existed in the static provider registry (as claude-sonnet-4-6
and claude-opus-4-7 do under vertex): resolveModelOrError had its own
ad-hoc resolution that never consulted the override, and even once
fixed, executeChatWithBreaker discarded the correctly-resolved format
before handleChatCore's own resolution ran a second time.

* docs: add changelog fragment for diegosouzapw#8909

* refactor(sse): extract shared Claude effort-model predicate

* fix(sse): strip Claude effort-suffix ids for any provider serving a real Claude model

* fix(sse): keep no-think and CC-discovery catalog variant roots unprefixed

* fix(dashboard): re-qualify no-think playground model ids correctly

* fix(sse): scope Vertex 404s to a per-model lockout via passthroughModels

* docs: add changelog fragment for the Claude catalog/dispatch fix

* fix(sse): align regex naming and changelog formatting

* fix(sse): clarify effort-variant strip comment and add cross-module drift guard

* fix(sse): disambiguate Vertex connection-wide vs per-model 403s

* docs: document Vertex 403 disambiguation in changelog fragment

* fix(sse): correlate reason and resource within the same ErrorInfo detail

* fix(sse): extract Vertex error classifier and rebaseline frozen file sizes

* test: register vertex-passthrough-model-lockout in stryker tap.testFiles

* fix(sse): reconciles rebase-onto-tip drift for 9006

Two categories of inherited base-branch breakage surfaced when
rebasing onto release/v3.8.50's latest tip, both confirmed unrelated
to this PR's own diff:

- check:file-size: base.ts and chat.ts drifted further past their
  frozen caps via already-merged commits (174e08b and others) that
  didn't rebaseline after growing them. Documented and bumped in
  file-size-baseline.json.
- chat-helpers.test.ts: two gpt-5.5 routing assertions predate diegosouzapw#9275
  (fix(routing): bare model ids route to codex first), which
  deliberately made gpt-5.5 route to codex unconditionally, regardless
  of which other providers are active. Confirmed via diegosouzapw#9275's own
  commit message and code comments this is intentional, not a
  regression; verified reproducible on the raw base tip alone, with
  no changes from this PR involved. Updated both assertions and their
  names to match the new, intentional default.

* ci: re-trigger checks after GitHub Actions incident (2026-08-07, resolved)

* ci: re-trigger checks (previous push event was dropped)

* fix(quality): rebaseline combo-routing-engine.test.ts own-comment growth

The ALL_ACCOUNTS_INACTIVE->ALL_TARGETS_SKIPPED fix (a32aed7) added explanatory comments (+7 lines), pushing the file past its frozen 3457 cap. CI's PR-mode check:file-size caught it; local check-file-size.mjs was not re-run after that specific commit.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant