Skip to content

fix: close stealth video error leaks - #3390

Merged
steebchen merged 10 commits into
theopenco:mainfrom
pacocartones:fix/videos-stealth-error-redaction
Aug 3, 2026
Merged

steebchen merged 10 commits into
theopenco:mainfrom
pacocartones:fix/videos-stealth-error-redaction

Conversation

@pacocartones

@pacocartones pacocartones commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor

Summary

The video pipeline forwards upstream provider error bodies to the client verbatim, skipping the stealth redaction every other pipeline already applies. fetchUpstreamJson (apps/gateway/src/videos/videos.ts:2788) throws HTTPException(status, { message: body.error.message }) on !response.ok — with the full raw text when the error isn't JSON — and re-forwards body.msg untouched on the application-error branch; app.ts's onError then renders that message to the client via renderGatewayError. Unlike chat, rerank and embeddings, videos.ts never calls shouldRedactProviderError.

For stealth providers this leaks exactly what #3340 and the stealth-provider-errors.ts invariant exist to protect: provider identity, hostnames and vendor markings inside the raw error body (e.g. quota exceeded at https://<secret-host> from avalanche, whose base URL is a deployment secret).

Fix

Mirror the sibling pattern (rerank.ts:1060, embeddings.ts:1347):

  • New pure helper videos/upstream-error.ts: clientFacingUpstreamMessage(providerId, statusCode, rawMessage) — redactedProviderErrorText(statusCode) when shouldRedactProviderError(providerId), the raw message unchanged otherwise. Same comment as the sibling sites.
  • fetchUpstreamJson takes an optional providerId and routes both throw branches through the helper (the Upstream provider error (<status>) fallback was already generic and stays as-is). Internal logger.warn keeps the full body, consistent with the rest of the codebase.
  • All 11 call sites pass providerContext.providerId (each already has it in scope). Non-stealth providers are byte-identical to before.

Tests

New pure spec videos/upstream-error.spec.ts (DB-light, no harness):

  • stealth (avalanche): the raw message never reaches the client — the response is exactly redactedProviderErrorText(500) and contains neither the secret host nor the original message;
  • non-stealth (openai): raw message passes through unchanged;
  • undefined provider: raw message passes through unchanged.

Adversarial cycle run: commenting the guard makes the stealth test fail (expected 'quota exceeded at https://plataforma-…' to be 'Upstream provider error (500…)'); restoring it passes again.

Verification

  • New spec 3/3 ✅ · normalize-streaming-error.spec.ts 13/13 ✅ (the stealth sibling suites stay green)
  • turbo run build --filter=gateway ✅ (10/10 tasks — full typecheck of videos.ts + helper + spec)
  • eslint ✅ · prettier ✅ (touched files)

Not verified: end-to-end against a real stealth provider (no DB/Docker on this machine) — the fetchUpstreamJson → HTTPException → renderGatewayError path is verified by reading, and the redaction decision itself by the new unit tests.


Disclosure: an AI coding assistant helped survey the pipeline and draft the test scaffolding. The diff is small and was verified the slow way: I read every call site, confirmed videos.ts is the only pipeline without the guard, and ran the revert-the-guard check above to prove the test actually catches the leak. I own the change and will follow up on review comments.

Summary by CodeRabbit

  • Bug Fixes

    • Sensitive upstream provider details are now hidden from error messages for supported providers.
    • Video creation, media uploads, and background processing return consistent generic errors when provider requests fail.
    • Non-sensitive provider errors continue to display their original details.
    • Failed video status responses now include the appropriate sanitized error message.
  • Tests

    • Added coverage for HTTP, application-level, and background video-processing error redaction.

Follow-up: async video-job redaction

This update also redacts stealth-provider failures before the worker persists the public job error, so the asynchronous status endpoint cannot expose upstream response text or network details.

It adds a full POST /v1/videos leak-regression case and a worker-to-status regression case, moves the HTTP helper into the common stealth-error module, and requires provider context at every video upstream request.

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@coderabbitai

coderabbitai Bot commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Warning

Review limit reached

@steebchen, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 16 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: caf30352-4afc-408d-92a2-bfdbc509cbfe

📥 Commits

Reviewing files that changed from the base of the PR and between b37ed91 and 2a231b0.

📒 Files selected for processing (2)
  • apps/gateway/src/videos/videos.spec.ts
  • apps/worker/src/services/video-jobs.ts

Walkthrough

Video gateway requests now redact stealth-provider HTTP and application errors. Worker persistence applies the same rule to video job failures. Provider identifiers flow through video creation and upload paths, with gateway and integration test coverage.

Changes

Video error redaction

Layer / File(s) Summary
Client-facing error policy
apps/gateway/src/lib/stealth-provider-errors.ts, apps/gateway/src/lib/stealth-provider-errors.spec.ts
Stealth-provider HTTP errors use status-based messages. Other providers retain raw upstream messages.
Gateway upstream handling and provider wiring
apps/gateway/src/videos/videos.ts, apps/gateway/src/stealth-error-redaction.spec.ts
fetchUpstreamJson formats HTTP and application-level errors and receives provider identifiers from video and upload operations.
Persisted video job errors
apps/worker/src/services/video-jobs.ts, apps/gateway/src/test-utils/mock-openai-server.ts, apps/gateway/src/videos/videos.spec.ts
Worker polling and finalization sanitize stealth-provider errors before persistence and logging. Failed job responses expose the sanitized message.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Sequence Diagram(s)

sequenceDiagram
  participant VideoClient
  participant Gateway
  participant UpstreamProvider
  participant Worker
  VideoClient->>Gateway: create video or upload media
  Gateway->>UpstreamProvider: send provider request
  UpstreamProvider-->>Gateway: HTTP or application error
  Gateway-->>VideoClient: formatted error message
  Worker->>UpstreamProvider: poll video job
  UpstreamProvider-->>Worker: failed job error
  Worker-->>VideoClient: persisted sanitized error
Loading

Possibly related PRs

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly identifies the main change: closing stealth-provider video error leaks.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
apps/gateway/src/videos/videos.ts (1)

2791-2791: 🔒 Security & Privacy | 🔵 Trivial | ⚡ Quick win

Make providerId required at this boundary.

All current video call sites pass providerContext.providerId, but providerId?: string allows a future call site to omit it and silently bypass stealth-provider redaction. If video requests always have a provider identifier, make the parameter required. Otherwise, document and test the intentional undefined-provider path.

Suggested type change
-	providerId?: string,
+	providerId: string,
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/gateway/src/videos/videos.ts` at line 2791, Make the providerId
parameter required in the affected video request boundary instead of optional,
using the surrounding function signature as the change point. Preserve the
existing providerContext.providerId call-site behavior and ensure the TypeScript
signature no longer permits undefined values.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@apps/gateway/src/videos/videos.ts`:
- Line 2791: Make the providerId parameter required in the affected video
request boundary instead of optional, using the surrounding function signature
as the change point. Preserve the existing providerContext.providerId call-site
behavior and ensure the TypeScript signature no longer permits undefined values.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: d69e3e8d-cdc4-4094-8b44-edc3857c690f

📥 Commits

Reviewing files that changed from the base of the PR and between 44da286 and beaaca5.

📒 Files selected for processing (3)
  • apps/gateway/src/videos/upstream-error.spec.ts
  • apps/gateway/src/videos/upstream-error.ts
  • apps/gateway/src/videos/videos.ts

@steebchen

Copy link
Copy Markdown
Member

Review

Verified the leak and the fix end-to-end against main.

fetchUpstreamJson (apps/gateway/src/videos/videos.ts:2786) threw HTTPException with the raw upstream error.message — and, when the body isn't JSON, with the whole response text wrapped into that field. app.ts:167-185 renders error.message verbatim via renderGatewayError. The bug is real and the fix is correct. Each link checks out:

  • avalanche declares env.required.baseUrl: "LLM_AVALANCHE_BASE_URL" (packages/models/src/providers.ts:593), so isStealthProvider("avalanche") is true.
  • Both the !response.ok branch and the body.msg/body.code application-error branch were unguarded; chat, embeddings (embeddings.ts:1347) and rerank (rerank.ts:1060) all guard.
  • All 11 fetchUpstreamJson call sites are updated; the reindentation is pure churn with no other change.
  • Post-fetch throws in the provider-specific creators are already generic ("Avalanche video response did not include a task id"), so the create path is fully closed.
  • Internal logger.warn keeps the full body — right call.

1. The async status path still leaks (biggest gap)

Video is asynchronous, so the most likely place a user sees a provider failure isn't POST /v1/videos — it's GET /v1/videos/:id. That path is untouched:

  • apps/worker/src/services/video-jobs.ts:2307-2316 (fetchGenericVideoStatus, and its avalanche siblings) throws new Error(body.error.message) with the raw upstream text.
  • On the poll-error limit that message is interpolated into the stored job error — `Video generation failed after ${n} consecutive polling errors: ${message}` → error: extractError(failedResponse) (line 2570). Line 2466 likewise copies extractError(enrichedUpstreamStatus) verbatim.
  • serializeVideoJob (videos.ts:2450) returns error: job.error ?? null straight to the client.

Network-level failures leak here too: an ECONNREFUSED/DNS error message embeds the upstream hostname, which for a stealth provider is the secret base URL — exactly what clientFacingUpstreamFailureMessage exists to prevent. Either extend the fix to the worker's terminal-error write, or say explicitly in the description that it's a follow-up; as written the PR reads as "video routes are now safe".

2. providerId?: string is fail-open

Every one of the 11 call sites already has providerContext.providerId in scope, so nothing needs the default. For a security invariant, make the parameter required — then a new call site can't silently reintroduce the leak and TS catches it.

3. The test can't catch the real regression

The new spec exercises a 4-line pure function; you could delete the third argument from createAvalancheVeoVideoJob and all three tests still pass. The repo already has the right harness: apps/gateway/src/stealth-error-redaction.spec.ts drives real routes (/v1/chat/completions, /v1/responses, /v1/messages, /v1/images/generations) against a leaky mock server and asserts the vendor marker never reaches the client. A /v1/videos case there is what actually guards this. The revert-the-guard check described in the description only proves the helper's if, not the wiring.

4. Helper is in the wrong place

apps/gateway/src/lib/stealth-provider-errors.ts is the single home for this family (clientFacingUpstreamFailureMessage, buildUpstreamErrorClientPayload, redactErrorDetails). A pipeline-local videos/upstream-error.ts fragments the invariant — someone reading the lib to enumerate redaction surfaces won't find it. Move it there, and since clientFacingUpstreamFailureMessage already covers network failures, name this one clientFacingUpstreamErrorMessage to keep the distinction legible.

5. Dead env manipulation in the spec

shouldRedactProviderError → isStealthProvider checks whether the provider definition declares an env.required.baseUrl key (providers.ts:1773-1781); it never reads process.env. So setting LLM_AVALANCHE_BASE_URL and the afterEach restore do nothing but imply a runtime dependency that doesn't exist. Drop both. The test's real premise — "avalanche is a stealth provider" — deserves an explicit comment instead, mirroring the pinned-mapping comment at stealth-error-redaction.spec.ts:33-37.

6. Nit

The fallback `Upstream provider error (${response.status})` and redactedProviderErrorText(status) (Upstream provider error (500 Internal Server Error)) are two shapes for one concept. Routing the fallback through redactedProviderErrorText unifies them.

Other checks

  • No regression elsewhere in the file: streamVideoFromUrl and the MiniMax retrieve path already use generic 502 messages; debugMode only writes llmgateway_raw_request/llmgateway_upstream_request into the DB record, never into the client response.
  • No perf impact — one string check per error path.
  • Non-stealth providers are byte-identical, as claimed.

Verdict

Direction is right and the change is low-risk. Before merge I'd want (2) the required param and (5) the dead env code — both one-liners. (3) route-level test and (4) helper placement are what I'd push for to match repo convention. (1) is the one that decides whether this closes the leak or half of it — worth an explicit call rather than shipping silently.

@pacocartones pacocartones changed the title fix(gateway): redact stealth provider upstream error bodies in video routes fix: close stealth video error leaks Aug 3, 2026
@pacocartones

Copy link
Copy Markdown
Contributor Author

Thanks for the thorough review — addressed in e76a8d5.

  • The worker now redacts stealth-provider terminal errors before persisting the public job error, for both provider-declared failures and polling/network failures. Raw diagnostic response data is no longer exposed through the serialized job error.
  • providerId is required by the video upstream helper, and all callers provide it.
  • The HTTP helper now lives in the common stealth-error module and routes the fallback through redactedProviderErrorText.
  • I removed the dead test environment setup and added a full POST /v1/videos regression with a deliberately leaky upstream mock, plus a worker-to-GET /v1/videos/:id regression for persisted terminal failures.

Validation: pre-commit hooks and targeted lint passed; stealth-provider-errors.spec.ts passes (13 tests); tsc --noEmit passes for both gateway and worker. I could not run the database-backed integration tests locally because PostgreSQL and Redis are not available in this environment; CI is now running them.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@apps/worker/src/services/video-jobs.ts`:
- Around line 582-593: Sanitize all persisted stealth-provider error data, not
only videoJob.error: update the upstreamStatusResponse writes in the relevant
job-update paths and the llmgateway_last_poll_error retry persistence to remove
provider-controlled message, URL, and provider identifier fields. Reuse
clientFacingVideoJobError or an equivalent centralized sanitizer, and add
assertions that persisted records contain none of those upstream values while
preserving unsanitized errors for non-stealth providers.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 6f829f70-9f9a-455b-9408-7974bde64d50

📥 Commits

Reviewing files that changed from the base of the PR and between beaaca5 and e76a8d5.

📒 Files selected for processing (7)
  • apps/gateway/src/lib/stealth-provider-errors.spec.ts
  • apps/gateway/src/lib/stealth-provider-errors.ts
  • apps/gateway/src/stealth-error-redaction.spec.ts
  • apps/gateway/src/test-utils/mock-openai-server.ts
  • apps/gateway/src/videos/videos.spec.ts
  • apps/gateway/src/videos/videos.ts
  • apps/worker/src/services/video-jobs.ts
🚧 Files skipped from review as they are similar to previous changes (1)
  • apps/gateway/src/videos/videos.ts

Comment on lines +582 to +593
function clientFacingVideoJobError(
providerId: string,
error: VideoJobRecord["error"],
): VideoJobRecord["error"] {
if (!error || !isStealthProvider(providerId as ProviderId)) {
return error;
}

return {
message: "Upstream provider error",
};
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟠 Major | 🏗️ Heavy lift

Sanitize all persisted stealth-provider error fields.

clientFacingVideoJobError only sanitizes videoJob.error. The same updates persist the raw upstream payload in upstreamStatusResponse at lines 2509-2513 and 2596. The retry path also persists the raw polling exception in llmgateway_last_poll_error at lines 2614-2618.

For stealth providers, remove or sanitize provider-controlled error fields before every upstreamStatusResponse write. Add assertions that the persisted record does not contain the upstream message, URL, or provider identifier.

Also applies to: 2481-2484, 2588-2591

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/worker/src/services/video-jobs.ts` around lines 582 - 593, Sanitize all
persisted stealth-provider error data, not only videoJob.error: update the
upstreamStatusResponse writes in the relevant job-update paths and the
llmgateway_last_poll_error retry persistence to remove provider-controlled
message, URL, and provider identifier fields. Reuse clientFacingVideoJobError or
an equivalent centralized sanitizer, and add assertions that persisted records
contain none of those upstream values while preserving unsanitized errors for
non-stealth providers.

@steebchen

Copy link
Copy Markdown
Member

Re-review at e76a8d5

Nice turnaround — five of the six points are properly addressed:

  • helper moved to lib/stealth-provider-errors.ts and renamed clientFacingUpstreamErrorMessage ✅
  • providerId is now required ✅
  • the dead process.env.LLM_AVALANCHE_BASE_URL scaffolding is gone ✅
  • the non-JSON/no-error.message fallback now routes through the helper too ✅
  • worker write sites redacted — and because deliverWebhook builds its payload from serializeVideoJob(job), redacting at write time also covers the callback_url webhook, which I had not called out ✅

Three things still need attention. I ran the specs locally at e76a8d5 (pnpm build:core first, then vitest run --no-file-parallelism) — 72 passed, 1 failed.

1. The new /v1/videos route test fails (blocking)

× /v1/videos hides the raw upstream error for a stealth provider
AssertionError: expected 'Internal Server Error' to be 'Upstream provider error (500 Internal…'
  Expected: "Upstream provider error (500 Internal Server Error)"
  Received: "Internal Server Error"

The test calls setupCreditsApiKey, which puts the project in credits mode — and resolveVideoProviderContext only reads the org's providerKey row in api-keys mode (videos.ts:1455-1470). In credits mode it reads LLM_AVALANCHE_* from the environment (videos.ts:1506+), and the suite's beforeAll only sets LLM_TUNDRA_*, LLM_GLACIER_* and LLM_OPENAI_*. So the request never reaches the leaky mock; it dies earlier and app.ts renders the generic 500. expectNoLeak passes vacuously.

Adding these two lines to the beforeAll env block (and to the savedEnv key list) makes it pass, and the stack trace then confirms it goes through the intended path — fetchUpstreamJson (videos.ts:2856) ← createAvalancheVeoVideoJob:

process.env.LLM_AVALANCHE_API_KEY = "avalanche-env-key";
process.env.LLM_AVALANCHE_BASE_URL = leakyServerUrl;

Switching the project to api-keys mode so the inserted providerKey is actually used would work too. The videos.spec.ts worker-path test passes as written, and the mock-openai-server.ts msg change doesn't disturb the other 46 video tests.

2. The raw stealth error still reaches the customer via log.upstreamResponse (blocking)

I extended the new worker test with a probe on the log row it produces:

log.errorDetails         = {"statusCode":502,"statusText":"failed","responseText":"Upstream provider error"}   ✅
log.internalErrorDetails = null
log.upstreamResponse     = {"error":{"details":{"msg":"SecretVendor error at https://api.secretvendor.com",…},
                             "message":"SecretVendor error at https://api.secretvendor.com"},…,
                            "avalanche_record_info":{"msg":"SecretVendor error at https://api.secretvendor.com",…}}

apps/api/src/routes/logs.ts:51 builds publicLogColumns by stripping only internalErrorDetails, so upstreamResponse is returned by the logs API and rendered in the dashboard. The video worker persists it unconditionally (video-jobs.ts:1913), which makes video the outlier — chat only stores upstreamResponse when x-debug is set (create-log-entry.ts:118-123). For a retain org that's the entire raw upstream body, plus llmgateway_last_poll_error, which carries the secret host verbatim on a network failure.

So the message-level leak is closed but the body-level one isn't. Either omit upstreamResponse from the log for stealth providers, or redact it on the way in. Keeping the raw copy on videoJob.upstreamStatusResponse (not customer-visible) preserves support's ability to debug.

3. Redacting at DB-write time has two side effects

Content-filter classification breaks for stealth providers. finalizeVideoJob derives the finish reason from the already redacted error (video-jobs.ts:1844-1849):

isContentFilterErrorText([jobToLog.error?.code, jobToLog.error?.message].filter(Boolean).join(" "))

With {message: "Upstream provider error"} and no code, a content-filtered avalanche job is now logged as upstream_error/UPSTREAM_ERROR instead of content_filter/CONTENT_FILTER — my probe confirms finishReason: upstream_error on the redacted row. Compute isContentFilterFailure from the raw error before redacting (or redact at serialization time rather than at write time).

No internal copy. insertLog deliberately keeps the raw error in internalErrorDetails (logs.ts:308-315); the video log insert bypasses insertLog and leaves that column null. Recoverable from videoJob.upstreamStatusResponse today, but it diverges from the documented invariant — worth mirroring while you're in there.

Nits

  • clientFacingVideoJobError returns {message: "Upstream provider error"}, dropping the status, while everything else uses redactedProviderErrorText(status) → Upstream provider error (500 Internal Server Error). The log row already stamps statusCode: 502, so redactedProviderErrorText(502) would keep one canonical string across both surfaces.
  • The worker re-implements the stealth predicate locally because lib/stealth-provider-errors.ts isn't importable across apps. Fine for two callers; if a third appears, the predicate and the redacted text belong in packages/shared.

@pacocartones

Copy link
Copy Markdown
Contributor Author

Thanks — excellent re-review. I addressed all three points in 3f0b6b50.

  1. The /v1/videos test now pins LLM_AVALANCHE_API_KEY and LLM_AVALANCHE_BASE_URL to the mock server for the suite, restoring the values afterwards. This makes the credits-mode path exercise the intended leaky upstream response instead of failing before the request.

  2. Video logs now have an explicit public/internal boundary for stealth providers: upstreamResponse is null in the customer-visible log, while the raw error is retained in internalErrorDetails. The raw polling status remains on videoJob.upstreamStatusResponse for worker-side diagnosis, but is not copied into the public log record.

  3. finalizeVideoJob derives content-filter classification from the raw upstream error before constructing public log fields. The public error is now the canonical Upstream provider error (502 Bad Gateway), while an extended regression test verifies all of the following together: no upstream host/provider text in the status response or public log, raw error retained internally, upstreamResponse absent publicly, and content_filter preserved as the finish reason.

CI for the final commit is running now: https://github.com/theopenco/llmgateway/actions/runs/30819873638

@pacocartones

Copy link
Copy Markdown
Contributor Author

Small test-only follow-up in 587e042: the regression now explicitly separates the internal row (where the raw error is intentionally retained) from its public projection before asserting that no secret is exposed. This is the final head under CI: https://github.com/theopenco/llmgateway/actions/runs/30820050420

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@apps/worker/src/services/video-jobs.ts`:
- Around line 1845-1850: Update the rawVideoError assignment near extractError
so it falls back to jobToLog.error when the upstreamStatusResponse is not an
object or extractError returns no error. Preserve the extracted upstream error
whenever it exists, ensuring finalized logs retain available error details.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: d2b78055-57de-4e1f-9ceb-a83bafbd07a4

📥 Commits

Reviewing files that changed from the base of the PR and between e76a8d5 and 587e042.

📒 Files selected for processing (2)
  • apps/gateway/src/videos/videos.spec.ts
  • apps/worker/src/services/video-jobs.ts
🚧 Files skipped from review as they are similar to previous changes (1)
  • apps/gateway/src/videos/videos.spec.ts

Comment on lines +1845 to +1850
const rawVideoError =
jobToLog.upstreamStatusResponse &&
typeof jobToLog.upstreamStatusResponse === "object" &&
!Array.isArray(jobToLog.upstreamStatusResponse)
? extractError(jobToLog.upstreamStatusResponse as Record<string, unknown>)
: jobToLog.error;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Fall back to jobToLog.error when the upstream payload has no error.

Lines 1845-1850 select null when upstreamStatusResponse is an object without an error field. This discards a non-null jobToLog.error. The finalized log then has no public or internal error details.

Use the extracted upstream error only when it exists. Otherwise, use jobToLog.error.

Proposed fix
-			const rawVideoError =
+			const upstreamVideoError =
 				jobToLog.upstreamStatusResponse &&
 				typeof jobToLog.upstreamStatusResponse === "object" &&
 				!Array.isArray(jobToLog.upstreamStatusResponse)
 					? extractError(jobToLog.upstreamStatusResponse as Record<string, unknown>)
-					: jobToLog.error;
+					: null;
+			const rawVideoError = upstreamVideoError ?? jobToLog.error;
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
const rawVideoError =
jobToLog.upstreamStatusResponse &&
typeof jobToLog.upstreamStatusResponse === "object" &&
!Array.isArray(jobToLog.upstreamStatusResponse)
? extractError(jobToLog.upstreamStatusResponse as Record<string, unknown>)
: jobToLog.error;
const upstreamVideoError =
jobToLog.upstreamStatusResponse &&
typeof jobToLog.upstreamStatusResponse === "object" &&
!Array.isArray(jobToLog.upstreamStatusResponse)
? extractError(jobToLog.upstreamStatusResponse as Record<string, unknown>)
: null;
const rawVideoError = upstreamVideoError ?? jobToLog.error;
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/worker/src/services/video-jobs.ts` around lines 1845 - 1850, Update the
rawVideoError assignment near extractError so it falls back to jobToLog.error
when the upstreamStatusResponse is not an object or extractError returns no
error. Preserve the extracted upstream error whenever it exists, ensuring
finalized logs retain available error details.

@steebchen

Copy link
Copy Markdown
Member

Re-review at b37ed91

All three round-2 findings are fixed, and I verified each one locally rather than by reading:

  • Route test — env vars added to the suite's beforeAll and the unused providerKey insert dropped. It now passes and genuinely reaches the leaky mock. ✅
  • log.upstreamResponse — nulled for stealth providers, with the assertion I was hoping for: expect(log!.upstreamResponse).toBeNull() plus a whole-public-row scan (const { internalErrorDetails: _, ...publicLog }) for both markers. That's a better boundary test than the one I suggested. ✅
  • Classification + internal copy — rawVideoError is derived from upstreamStatusResponse before redaction, so isContentFilterErrorText sees the real text again, and internalErrorDetails now carries the raw payload the way insertLog does. The test pins finishReason === "content_filter" alongside the redacted public errorDetails, which locks the regression out. ✅

Verification: videos.spec.ts + stealth-provider-errors.spec.ts + stealth-error-redaction.spec.ts → 73 passed, 0 failed. turbo run build --filter=gateway and --filter=worker clean, eslint clean on the touched files.

One trap worth knowing about, since it cost me a false alarm: videos.spec.ts imports processPendingVideoJobs from the built worker package, not from source. With a stale worker/dist the new content_filter assertion fails with expected 'upstream_error' to be 'content_filter' even though the code is correct. Rebuild the worker (or run pnpm build:core) before running that spec.

One thing to fix before merge

pnpm format hasn't been run — prettier --check fails on both changed files:

[warn] apps/worker/src/services/video-jobs.ts
[warn] apps/gateway/src/videos/videos.spec.ts

Three hunks: the extractError( continuation indent and the ${\n VIDEO_JOB_PUBLIC_ERROR_STATUS_CODE\n} interpolation in video-jobs.ts, and two long lines in the spec. pnpm format fixes all of them.

Nits, take or leave

  • The redacted string is now built inline in two places with the same template (clientFacingVideoJobError and publicErrorDetails). A single videoJobRedactedErrorText() next to the constants would keep them from drifting.
  • VIDEO_JOB_PUBLIC_ERROR_STATUS_CODE = 502 is synthetic — the status poll returns HTTP 200 with a failed job payload, so there is no real upstream status here, unlike redactedProviderErrorText(status) elsewhere. It matches the statusCode: 502 this code already stamped, so it's the right call; a one-line comment saying the 502 is a stand-in would stop a future reader from "fixing" it.
  • Redaction now also swallows gateway-authored messages: a stealth job that hits buildVideoJobTimeoutResponse ("Video generation timed out after N seconds without reaching a terminal state", code: "timeout") surfaces as Upstream provider error (502 Bad Gateway). That text contains nothing provider-identifying, so users lose a useful signal for free. Scoping the redaction to provider-sourced errors would keep it — though note the poll-error-limit message does embed raw upstream text, so only the pure timeout case is safely exempt.
  • rawErrorDetails.statusText is jobToLog.status ("failed") while the public one is "Bad Gateway". Harmless, just two shapes for the same row.
  • I checked the remaining debug payloads: rawRequest / upstreamRequest are still stored for stealth providers, but they only carry our own request shape and the catalogue model id — no host, no vendor marker — so there's nothing left to redact there.

LGTM once pnpm format lands.

steebchen and others added 2 commits August 3, 2026 17:28
Run pnpm format over the changed files and hoist the redacted
error text into a single constant, with a note on why the public
status code is a hard-coded 502.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@steebchen
steebchen enabled auto-merge August 3, 2026 16:32
@steebchen
steebchen added this pull request to the merge queue Aug 3, 2026
Merged via the queue into theopenco:main with commit 286b02b Aug 3, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants