Skip to content

fix(api): record audio transcription/translation/speech requests in call_logs (#13544) - #13803

Merged
diegosouzapw merged 2 commits into
release/v3.8.51from
fix/13544-audio-transcriptions-not-logged
Sep 16, 2026
Merged

diegosouzapw merged 2 commits into
release/v3.8.51from
fix/13544-audio-transcriptions-not-logged

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

Closes #13544

Root cause

open-sse/handlers/audioTranscription.ts and src/app/api/v1/audio/transcriptions/route.ts
never called saveCallLog() (@/lib/usageDb), unlike every other proxied API surface
(embeddings, images, video, rerank, search). Successful and failed
/v1/audio/transcriptions requests were therefore invisible in Dashboard → Request Logs,
even though the upstream provider was dispatched correctly and the transcription was
returned to the client. The reporter's secondary point was also confirmed: the response
route hardcodes costUsd: 0 with a comment claiming "no text body available" — stale for
the OpenAI-compatible default handler, which already buffers the full upstream JSON body
(including a provider's usage:{type:"duration",seconds:N}) before returning it; it was
just never parsed for logging.

The identical gap exists in the sibling /v1/audio/translations and /v1/audio/speech
routes — same code shape, not user-reported, fixed alongside since it's the same class of
defect.

Fix

  • Threaded policy.apiKeyInfo.{id,name} from enforceApiKeyPolicy() into
    transcribeWithModel() / translateWithModel() (and the postHandler closure for
    speech), matching how rerank.ts and embeddings.ts already do this.
  • Added saveCallLog() on both the success and error branch of all three routes, with
    provider, model, connectionId, duration, apiKeyId, apiKeyName — mirroring the
    canonical shape in open-sse/handlers/embeddings.ts / src/app/api/v1/rerank/route.ts.
  • Added peekDurationUsage() in the transcriptions route: a best-effort peek (via
    response.clone().json(), never touching the original stream) at a provider's
    usage:{type:"duration",seconds:N} body, persisted on the call_logs row's
    responseBody for audit ahead of a future per-second cost rule. No cost pricing engine
    is added here — costUsd stays 0 for transcription, matching the existing comment's
    actual (if previously mis-explained) behavior.
  • Combo fan-out (handleComboChat → transcribeWithModel/translateWithModel per target)
    now receives the caller's apiKeyId/apiKeyName too, so combo-routed audio requests are
    attributed correctly as well.

Regression test

tests/unit/issue-13544-audio-transcription-call-log.test.ts — drives a real
POST /v1/audio/transcriptions through the actual route (provider-node resolution,
enforceApiKeyPolicy, upstream dispatch) against a mocked OpenAI-compatible provider.

RED (on unfixed release/v3.8.51 tip):

✖ #13544: a successful transcription through an OpenAI-compatible provider node creates a call_logs entry
  AssertionError [ERR_ASSERTION]: expected a call_logs row for the successful /v1/audio/transcriptions request
  (none was created — the transcription path bypasses the normal call-log pipeline, #13544)

GREEN (after the fix):

✔ #13544: a successful transcription through an OpenAI-compatible provider node creates a call_logs entry (32359ms)
✔ #13544: a failed upstream transcription request also creates a call_logs entry (410ms)
ℹ tests 2
ℹ pass 2
ℹ fail 0

The success-path assertions also confirm the row carries provider/model/status/
apiKeyId, and that getCallLogById() returns the persisted usage:{type:"duration", seconds:3} on responseBody (issue Validation Plan steps 4 and 5).

Gates run

  • npx eslint --suppressions-location config/quality/eslint-suppressions.json <changed files> → 0 errors/warnings
  • npm run typecheck:core → 0 errors
  • node scripts/check/check-test-discovery.mjs → OK, 0 new orphans
  • node scripts/check/check-cognitive-complexity.mjs → OK (1276 violations vs baseline 1437 — improved, not regressed)
  • node scripts/check/check-file-size.mjs → no ✗ on any file this PR touches
  • node scripts/check/check-complexity.mjs → did not finish within the session's available
    time due to heavy concurrent load from several other active sessions on the shared devbox
    running the same repo-wide scan in parallel (confirmed via ps); this PR's diff is a
    small, flat threading of a few extra parameters plus straight-line if/else saveCallLog
    calls copied from the already-passing rerank.ts/embeddings.ts pattern, so it is very
    unlikely to move this ratchet. Flagging for CI to confirm rather than asserting a result
    I couldn't observe locally.
  • Full regression + sibling test files (tests/unit/issue-13544-audio-transcription-call-log.test.ts,
    audio-transcriptions-combo-resolution.test.ts, audio-transcription-handler.test.ts,
    audio-transcription-openrouter.test.ts, audio-transcription-opus-filename.test.ts,
    audio-translations-combo-resolution.test.ts, audio-translations-route.test.ts,
    audio-speech-dynamic-node-9096.test.ts, audio-speech-fishaudio.test.ts,
    audio-speech-handler.test.ts, audio-speech-ogg-alias-10587.test.ts) → 88/88 pass, no
    regressions.

Existing tests aligned

None — no existing assertion encoded the old buggy contract; all pre-existing audio
test files pass unmodified.

Not covered here

  • No new per-second cost/pricing rule for transcription duration (out of scope per the
    triage plan — costUsd stays 0, unchanged behavior; only the upstream usage value is
    now persisted for future consumption).
  • Non-default transcription provider formats (Deepgram, AssemblyAI, Gladia, etc.) get the
    core saveCallLog fix but not the duration-usage enrichment, since only the default
    OpenAI-compatible branch exposes a buffered JSON body to peek at.

diegosouzapw and others added 2 commits September 15, 2026 18:16
…all_logs (#13544)

open-sse/handlers/audioTranscription.ts and src/app/api/v1/audio/transcriptions/route.ts
never called saveCallLog(), unlike every other proxied endpoint (embeddings, images,
video, rerank, search). Successful and failed /v1/audio/transcriptions requests were
therefore invisible in Dashboard -> Request Logs. Threaded the API key metadata from
enforceApiKeyPolicy() into the handler and added the missing saveCallLog() call on both
the success and error path, plus a best-effort peek at upstream `usage:{type:"duration"}`
so it is persisted for future cost-pipeline consumption. Applied the identical fix to the
sibling /v1/audio/translations and /v1/audio/speech routes, which had the same gap.

Regression test: tests/unit/issue-13544-audio-transcription-call-log.test.ts
@diegosouzapw
diegosouzapw merged commit d7d518a into release/v3.8.51 Sep 16, 2026
19 of 21 checks passed
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…all_logs (diegosouzapw#13544) (diegosouzapw#13803)

Merged in the 2026-09-16 sweep of the maintainer's own open PRs, at the owner's explicit instruction. No push was made to the PR branch: the merge took the head as the owning session left it (verified OPEN, non-draft and MERGEABLE against the release tip immediately before merging).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] audio: successful /v1/audio/transcriptions requests are not recorded in call_logs or proxy_logs

1 participant