Skip to content

fix(resilience): winning failover attempt's call log dropped (call_logs.id UNIQUE clash) #13481

Description

@Xore

OmniRoute Version

3.8.51 (built from source, release/v3.8.51)

Installation Method

Built from source

Operating System

Linux

Node.js Version

22.23.2

Provider(s) Involved

Any — reproduced with opencode (no-auth) as the failing step and openrouter as the succeeding step of a priority combo

Model(s) Involved

opencode/deepseek-v4-flash-free (fails with 400), openrouter/nvidia/nemotron-3-ultra-550b-a55b:free (succeeds)

Client Tool

Any OpenAI-compatible client (POST /v1/chat/completions, streaming)

Description

When a priority combo fails over from one member to the next inside a single client request, the winning attempt's call log is never persisted. Every attempt of the request is saved with id: pendingRequestId (open-sse/handlers/chatCore/attemptLogging.ts, saveCallLog({ id: pendingRequestId, ... })), and pendingRequestId is allocated once per client request in open-sse/handlers/chatCore.ts (trackPendingRequest(...)). The first attempt that reaches persistAttemptLogs (the failed member) inserts the row; the successful member's row then hits the primary key and is dropped:

[callLogs] Failed to save call log: SqliteError: UNIQUE constraint failed: call_logs.id

Net effect: for every request that needed a failover, call_logs (dashboard "Logs", usage/cost accounting, combo_step_id stats) contains only the failed step(s) and never the 200 that actually answered the client. A combo whose first healthy member sits behind a dead one looks 100% broken in the dashboard while clients are being served normally. Combos that succeed on the first member are unaffected, which is why this is easy to miss.

Steps to Reproduce

  1. Create a priority combo with two members: first a model that fails deterministically on the provider side (here opencode/deepseek-v4-flash-free, which currently answers 400 "Model is unavailable"), second any working model (here openrouter/nvidia/nemotron-3-ultra-550b-a55b:free).
  2. Send one streaming POST /v1/chat/completions with "model": "<combo>".
  3. The client receives a complete answer from member 2.
  4. Open Logs in the dashboard (or SELECT status, model, combo_step_id FROM call_logs WHERE combo_name='<combo>').

Expected Behavior

Two rows for the request: the 400 from member 1 and the 200 from member 2 (each attempt with its own id, sharing correlation_id / combo_execution_key), or at minimum the terminal successful attempt.

Actual Behavior

Only the 400 row exists. The application log shows the collision right after the stream completes:

23:41:56.721Z [ERROR] [400]: Error from provider (Console): Upstream request failed: Model is unavailable.
23:41:56.837Z COMBO   Trying model openrouter/nvidia/nemotron-3-ultra-550b-a55b:free
23:42:38.994Z [STREAM] OPENROUTER | nvidia/nemotron-3-ultra-550b-a55b:free | 41064ms | complete
23:42:38.893Z [USAGE] OPENROUTER | in=56949 | out=133 | reasoning=48
23:42:39.088Z callLogs [callLogs] Failed to save call log: SqliteError: UNIQUE constraint failed: call_logs.id
sqlite> select timestamp,status,combo_step_id,tokens_in,tokens_out from call_logs where combo_name='<combo>' order by timestamp;
2026-09-12T23:41:56.831Z|400|<combo>-model-13-oc-deepseek-v4-flash-free|0|0
2026-09-12T23:42:44.344Z|400|<combo>-model-13-oc-deepseek-v4-flash-free|0|0
2026-09-12T23:43:02.703Z|400|<combo>-model-13-oc-deepseek-v4-flash-free|0|0

(7 completed openrouter streams in the same window, 0 rows for them; 11 UNIQUE constraint failed lines since the last restart, 29 in the previous log file.)

Test Impact

Needs a new unit test

Error Logs / Output

{"timestamp":"2026-09-12T23:42:39.088Z","level":"error","component":"callLogs","message":"[callLogs] Failed to save call log: SqliteError: UNIQUE constraint failed: call_logs.id"}

Additional Context

  • Not the same as fix(resilience): image/video combo sequential execution blocks healthy providers; call_logs UNIQUE collisions #13099: that one is generateLogId() colliding across two processes sharing a DATA_DIR. Here a single process supplies a non-empty id (pendingRequestId) for every attempt of one request, so generateLogId() is never involved.
  • src/lib/usage/callLogs.ts only falls back to generateLogId() when entry.id is empty; a non-empty caller-supplied id is trusted as unique.
  • persistAttemptLogs is invoked from ~10 sites in chatCore.ts (per-attempt error paths and the final success path); all of them pass the same pendingRequestId.
  • Suggested direction: give each attempt its own log id (e.g. the per-attempt traceId that attemptLogging.ts already documents as "MUST match the id emitted in request.started"), keep pendingRequestId in correlation_id; or make saveCallLog upsert on conflict so the terminal attempt wins. Either way the successful attempt must never be silently dropped.

Validation Plan

  • New unit test for persistAttemptLogs/saveCallLog: two attempts under one pendingRequestId (400 then 200) must both be readable via getCallLogById / the logs listing.
  • node --import tsx/esm --test tests/unit/callLogs.test.ts (and the combo failover tests under tests/unit/ that cover comboAttemptLoop).
  • Manual: repeat steps above, confirm the 200 row appears with the correct combo_step_id and token counts.

Activity

  1. changed the title [-][BUG] Combo failover: successful attempt's call log dropped (UNIQUE constraint failed: call_logs.id — all attempts share pendingRequestId)[/-] [+]fix(resilience): winning failover attempt's call log dropped (call_logs.id UNIQUE clash)[/+] on Sep 14, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions