Conversation
WalkthroughTreat 4xx responses (including 404) as retryable for auto-selected providers, add a one-shot 404 trigger in the mock OpenAI test server, and add tests verifying fallback retries succeed when an initial provider returns 404. Changes
Sequence Diagram(s)sequenceDiagram
participant Client as Client
participant Gateway as Gateway
participant ProviderA as Provider A
participant ProviderB as Provider B
participant DB as DB
Client->>Gateway: send request (no explicit provider)
Gateway->>Gateway: route -> auto-select Provider A
Gateway->>ProviderA: forward request
ProviderA-->>Gateway: 404 (model-not-found)
Gateway->>DB: log failed routing (status 404, retried=true)
Gateway->>ProviderB: retry request to fallback provider
ProviderB-->>Gateway: 200 OK (response)
Gateway->>DB: log successful routing (link retry -> success)
Gateway-->>Client: return successful response
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related PRs
Suggested labels
Suggested reviewers
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Pull request overview
Adds 404 as a retryable status code in the gateway's retry-with-fallback logic, so that auto-routed requests that hit a provider returning 404 (e.g., model-not-found or access drift) fall back to the next scored provider.
Changes:
isRetryableErrornow treats 404 as retryable alongside 429, 5xx, and network errors- New unit tests for 404 retry behavior and an integration test exercising 404→fallback success
- Mock OpenAI server extended with a
TRIGGER_STATUS_404_ONCEone-shot trigger
Reviewed changes
Copilot reviewed 4 out of 4 changed files in this pull request and generated no comments.
| File | Description |
|---|---|
apps/gateway/src/chat/tools/retry-with-fallback.ts |
Add 404 to retryable status codes |
apps/gateway/src/chat/tools/retry-with-fallback.spec.ts |
Unit tests for 404 retry behavior |
apps/gateway/src/fallback.spec.ts |
Integration test for 404 fallback |
apps/gateway/src/test-utils/mock-openai-server.ts |
One-shot 404 trigger for tests |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
You can also share your feedback on Copilot code review. Take the survey.
There was a problem hiding this comment.
Actionable comments posted: 1
🧹 Nitpick comments (1)
apps/gateway/src/fallback.spec.ts (1)
830-875: Assert that the provider actually changes.Because the mock
TRIGGER_STATUS_404_ONCEsucceeds on the second matching request regardless of provider, this spec currently proves “retry” but not “fallback”. Please compare the first and lastrouting[*].providervalues (and/orfailedLog.usedProvidervssuccessLog.usedProvider) so a same-provider retry cannot pass.💡 Suggested tightening
expect(json.metadata.routing[0]).toHaveProperty("status_code", 404); expect(json.metadata.routing[0]).toHaveProperty("succeeded", false); - expect( - json.metadata.routing[json.metadata.routing.length - 1], - ).toHaveProperty("succeeded", true); + const lastAttempt = + json.metadata.routing[json.metadata.routing.length - 1]; + expect(lastAttempt).toHaveProperty("succeeded", true); + expect(json.metadata.routing[0].provider).not.toBe(lastAttempt.provider); @@ const failedLog = logs.find((l: Log) => l.hasError); expect(failedLog).toBeDefined(); + expect(failedLog!.usedProvider).not.toBe(successLog!.usedProvider); expect(failedLog!.finishReason).toBe("upstream_error"); expect(failedLog!.retried).toBe(true);🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@apps/gateway/src/fallback.spec.ts` around lines 830 - 875, Update the test "non-streaming: retries on 404 and succeeds on fallback provider" to assert that the request actually fell back to a different provider: after parsing json.metadata.routing, compare the first routing entry's provider to the last routing entry's provider (e.g., json.metadata.routing[0].provider !== json.metadata.routing[json.metadata.routing.length - 1].provider) and/or compare failedLog.usedProvider to successLog.usedProvider (they must differ) so a same-provider retry cannot satisfy the test.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@apps/gateway/src/test-utils/mock-openai-server.ts`:
- Around line 241-254: The generic status parser (extractStatusCodeTrigger) is
matching the prefix TRIGGER_STATUS_404 inside TRIGGER_STATUS_404_ONCE so the
earlier handler returns before the "_ONCE" branch (which uses fail404OnceCounter
and userMessage) can run; update the mock-openai-server logic to check for the
specific *_ONCE triggers before invoking the generic status parser (or update
extractStatusCodeTrigger to prefer full-string matches including the "_ONCE"
suffix), ensuring TRIGGER_STATUS_404_ONCE is detected first so the
fail404OnceCounter branch executes exactly once and subsequent requests follow
the normal flow.
---
Nitpick comments:
In `@apps/gateway/src/fallback.spec.ts`:
- Around line 830-875: Update the test "non-streaming: retries on 404 and
succeeds on fallback provider" to assert that the request actually fell back to
a different provider: after parsing json.metadata.routing, compare the first
routing entry's provider to the last routing entry's provider (e.g.,
json.metadata.routing[0].provider !==
json.metadata.routing[json.metadata.routing.length - 1].provider) and/or compare
failedLog.usedProvider to successLog.usedProvider (they must differ) so a
same-provider retry cannot satisfy the test.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: f91d1a8f-f020-4330-9e3e-fe2bf3211a0a
📒 Files selected for processing (4)
apps/gateway/src/chat/tools/retry-with-fallback.spec.tsapps/gateway/src/chat/tools/retry-with-fallback.tsapps/gateway/src/fallback.spec.tsapps/gateway/src/test-utils/mock-openai-server.ts
| if (userMessage.includes("TRIGGER_STATUS_404_ONCE")) { | ||
| fail404OnceCounter++; | ||
| if (fail404OnceCounter === 1) { | ||
| c.status(404); | ||
| return c.json({ | ||
| error: { | ||
| message: "The model 'nonexistent-model' does not exist.", | ||
| type: "invalid_request_error", | ||
| param: "model", | ||
| code: "model_not_found", | ||
| }, | ||
| }); | ||
| } | ||
| } |
There was a problem hiding this comment.
The _ONCE trigger is currently swallowed by the generic status parser.
extractStatusCodeTrigger() on Line 60 already matches the TRIGGER_STATUS_404 prefix inside TRIGGER_STATUS_404_ONCE, and the earlier generic handler starting at Line 217 returns before this branch is reached. In apps/gateway/src/fallback.spec.ts Line 842, that means every attempt sees a 404 instead of “404 once, then succeed”.
💡 Suggested fix
function extractStatusCodeTrigger(
content: string,
): { statusCode: number; errorResponse: object } | null {
- const match = content.match(/TRIGGER_STATUS_(\d{3})/);
+ const match = content.match(/\bTRIGGER_STATUS_(\d{3})(?!_ONCE)\b/);
if (!match) {
return null;
}🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@apps/gateway/src/test-utils/mock-openai-server.ts` around lines 241 - 254,
The generic status parser (extractStatusCodeTrigger) is matching the prefix
TRIGGER_STATUS_404 inside TRIGGER_STATUS_404_ONCE so the earlier handler returns
before the "_ONCE" branch (which uses fail404OnceCounter and userMessage) can
run; update the mock-openai-server logic to check for the specific *_ONCE
triggers before invoking the generic status parser (or update
extractStatusCodeTrigger to prefer full-string matches including the "_ONCE"
suffix), ensuring TRIGGER_STATUS_404_ONCE is detected first so the
fail404OnceCounter branch executes exactly once and subsequent requests follow
the normal flow.
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
apps/gateway/src/chat/tools/get-finish-reason-from-error.ts (1)
32-48:⚠️ Potential issue | 🔴 CriticalDead code: xAI content filter check for 403 is unreachable.
The 401/403 check at lines 33-35 returns
"gateway_error"before the xAI content filter check at lines 43-48 can execute. This means xAI Grok responses with"Content violates usage guidelines"will incorrectly return"gateway_error"instead of"content_filter".🐛 Proposed fix: Move content filter check before 401/403 handling
// 404 from upstream provider indicates model/endpoint not found at provider if (statusCode === 404) { return "upstream_error"; } + // xAI (Grok) content safety violations (e.g. SAFETY_CHECK_TYPE_CSAM, usage guidelines) + if ( + statusCode === 403 && + errorText?.includes("Content violates usage guidelines") + ) { + return "content_filter"; + } + // 401/403 indicate gateway-side auth/configuration issues if (statusCode === 401 || statusCode === 403) { return "gateway_error"; } // Azure OpenAI content filter (ResponsibleAIPolicyViolation) if (errorText?.includes("ResponsibleAIPolicyViolation")) { return "content_filter"; } - // xAI (Grok) content safety violations (e.g. SAFETY_CHECK_TYPE_CSAM, usage guidelines) - if ( - statusCode === 403 && - errorText?.includes("Content violates usage guidelines") - ) { - return "content_filter"; - }🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@apps/gateway/src/chat/tools/get-finish-reason-from-error.ts` around lines 32 - 48, The xAI/Grok content-filter branch is unreachable because the earlier 401/403 check returns "gateway_error" before the errorText check runs; in get-finish-reason-from-error (the function handling statusCode and errorText) move the xAI content safety check that looks for "Content violates usage guidelines" (and the Azure "ResponsibleAIPolicyViolation" check) to run before the generic 401/403 block, or adjust the conditional ordering so that when errorText includes the content-filter strings it returns "content_filter" even if statusCode === 403; update the logic in that function accordingly to ensure content_filter takes precedence over gateway_error for those errorText matches.
🧹 Nitpick comments (1)
apps/gateway/src/chat/tools/get-finish-reason-from-error.spec.ts (1)
73-76: Consider adding test coverage for xAI content filter edge case.There's an existing test for 401/403 returning
"gateway_error", but no test verifying the xAI content filter case where a 403 with"Content violates usage guidelines"should return"content_filter". This test would help catch the ordering bug mentioned in the implementation review.🧪 Suggested test to add
it("returns content_filter for xAI 403 content violation before gateway_error check", () => { expect( getFinishReasonFromError(403, "Content violates usage guidelines"), ).toBe("content_filter"); });🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@apps/gateway/src/chat/tools/get-finish-reason-from-error.spec.ts` around lines 73 - 76, Add a unit test in get-finish-reason-from-error.spec.ts to cover the xAI content filter edge case: call getFinishReasonFromError with status 403 and the message "Content violates usage guidelines" and assert it returns "content_filter" (this ensures the xAI content-filter branch in getFinishReasonFromError is checked before the generic 403 -> "gateway_error" handling).
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@apps/gateway/src/chat/tools/get-finish-reason-from-error.ts`:
- Around line 32-48: The xAI/Grok content-filter branch is unreachable because
the earlier 401/403 check returns "gateway_error" before the errorText check
runs; in get-finish-reason-from-error (the function handling statusCode and
errorText) move the xAI content safety check that looks for "Content violates
usage guidelines" (and the Azure "ResponsibleAIPolicyViolation" check) to run
before the generic 401/403 block, or adjust the conditional ordering so that
when errorText includes the content-filter strings it returns "content_filter"
even if statusCode === 403; update the logic in that function accordingly to
ensure content_filter takes precedence over gateway_error for those errorText
matches.
---
Nitpick comments:
In `@apps/gateway/src/chat/tools/get-finish-reason-from-error.spec.ts`:
- Around line 73-76: Add a unit test in get-finish-reason-from-error.spec.ts to
cover the xAI content filter edge case: call getFinishReasonFromError with
status 403 and the message "Content violates usage guidelines" and assert it
returns "content_filter" (this ensures the xAI content-filter branch in
getFinishReasonFromError is checked before the generic 403 -> "gateway_error"
handling).
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: 81e27dc6-c3d4-4f93-807b-471176d473bb
📒 Files selected for processing (4)
apps/gateway/src/chat/tools/get-finish-reason-from-error.spec.tsapps/gateway/src/chat/tools/get-finish-reason-from-error.tsapps/gateway/src/chat/tools/retry-with-fallback.spec.tsapps/gateway/src/chat/tools/retry-with-fallback.ts
Summary
4xxprovider responses so provider-specific client and access errors can fall through to the next scored provider immediately401,403,404, and429stay in the gateway/upstream buckets and other4xxresponses becomeclient_error404trigger for the fallback specTest plan
pnpm vitest run apps/gateway/src/chat/tools/retry-with-fallback.spec.tspnpm vitest run apps/gateway/src/chat/tools/get-finish-reason-from-error.spec.ts apps/gateway/src/chat/tools/retry-with-fallback.spec.tspnpm exec prettier --check apps/gateway/src/chat/tools/get-finish-reason-from-error.ts apps/gateway/src/chat/tools/get-finish-reason-from-error.spec.ts apps/gateway/src/chat/tools/retry-with-fallback.ts apps/gateway/src/chat/tools/retry-with-fallback.spec.ts apps/gateway/src/test-utils/mock-openai-server.ts apps/gateway/src/fallback.spec.tspnpm vitest run apps/gateway/src/fallback.spec.ts(blocked in this environment: local PostgreSQL and Redis are not running)🤖 Generated with Codex
Summary by CodeRabbit
Bug Fixes
Tests