Skip to content

fix(logging): widen the trace id so it can key call_logs - #14341

Closed
abhisheksharma2411 wants to merge 1 commit into
diegosouzapw:release/v3.8.51from
abhisheksharma2411:fix/call-log-id-collision-14338
Closed

abhisheksharma2411 wants to merge 1 commit into
diegosouzapw:release/v3.8.51from
abhisheksharma2411:fix/call-log-id-collision-14338

Conversation

@abhisheksharma2411

Copy link
Copy Markdown
Contributor

Fixes #14338.

The id is 24 bits, and it is a primary key

call_logs.id is id TEXT PRIMARY KEY. The value written to it is the trace id chatCore mints for each attempt:

// open-sse/handlers/chatCore.ts
const traceId = globalThis.crypto.randomUUID().slice(0, 6);

Six hex characters — 16,777,216 values. That id was introduced as a log-correlation token (its own comment said so: "this id is a log-correlation token, not a security secret"), and #13481 later made it the call_logs row key so each combo attempt gets its own row and pairs with its request.started dashboard event. The width was never revisited for the second job.

It is not a concurrency race

The report's suggested repro — fire ~50 parallel requests — is the wrong experiment, and worth correcting because it makes the bug look load-dependent when it isn't. Two inserts in the same millisecond are not required. The collision is birthday-shaped against every row already stored: with N rows in the table, each new insert collides with probability N / 2²⁴.

rows in call_logs P(collision) per insert expected collisions per 1,000 inserts
50,000 0.298% 3.0
100,000 0.596% 6.0
268,000 1.597% 16.0
500,000 2.980% 29.8

That is exactly the reported shape — 8 failures in a 1-hour window under normal workload on a container whose table holds a few hundred thousand rows — and it explains why it reads as "periodic bursts": the bursts are just where the requests are, not where the contention is. It also means the rate grows as the table fills and drops after rotation, which no amount of insert serialisation would change.

Why the row vanishes silently

The losing insert throws inside saveCallLogOperation, and the catch logs and swallows:

} catch (error) {
  console.error("[callLogs] Failed to save call log:", ...);
}

saveCallLog still resolves, so nothing upstream can tell. The chat completion succeeds; only the analytics row is gone — which is why it surfaces as under-counted combo provider stats and cost rollups rather than as an error.

Reproduced first

The new test drives the real save path and shows the second write lost with saveCallLog still resolving happily:

[callLogs] Failed to save call log: SqliteError: UNIQUE constraint failed: call_logs.id
✔ a duplicate call-log id is dropped silently, not surfaced

— the same line as the container log in the report.

Fix

Move the generator into open-sse/handlers/chatCore/traceId.ts, so the constraint is written down where the value is produced rather than inferred from a call site 500 lines away, and widen it to 16 hex chars (64 bits):

export const TRACE_ID_HEX_CHARS = 16;

export function createTraceId(): string {
  return globalThis.crypto.randomUUID().replace(/-/g, "").slice(0, TRACE_ID_HEX_CHARS);
}

16 hex chars matches the width already used in this codebase — randomHex(16) in grok-web.ts, randomId(16) in the OTEL exporter — so [STAGE_TRACE] lines stay readable, and it takes the 268k-row collision probability from 1.6% to about 1.5 × 10⁻¹⁴ per insert.

Deliberately not changed: the id remains a single value shared by call_logs.id, request.started and request.completed. resolveRequestLifecycleEvent's contract is that the row id mirrors the paired event's traceId, so giving the row its own key would have broken dashboard pairing to fix an entropy problem.

Checks

tests/unit/call-log-id-collision-14338.test.ts 3 pass / 0 fail
chatcore + call-log + attempt-logging suites 735 tests, 734 pass / 1 fail
typecheck (tsc --noEmit) clean on the touched files
eslint, prettier clean

The single failure is call-log-file-rotation.test.ts → "orphan cleanup scans at most 100 candidates". I stashed the change and re-ran it on an otherwise clean release/v3.8.51 — it fails identically there, so it is a pre-existing base red and not this change.

Mutation-tested 2/2:

mutation result
TRACE_ID_HEX_CHARS back to 6 (the original id) ✗ width assertion, and ✗ a real duplicate (0a7381) inside 200k draws
drop .replace(/-/g, "") so a dash lands in the id ✗ charset assertion

The 200k-draw test is sized so it fails ~certainly on a 24-bit id (P ≈ 1 by 50k draws) while staying non-flaky on 64 bits (P ≈ 1e-9 across the whole run).

One thing I did not fix

The swallow in saveCallLogOperation means any future persistence failure is invisible the same way — a full disk, a locked DB, a schema drift. Widening the id removes the cause we know about, not the blind spot. Happy to follow up with a counter or a throttled warn if you'd like that separated out; it changes error-handling behaviour, so it didn't belong in a fix aimed at the collision itself.

`call_logs.id` is `id TEXT PRIMARY KEY` and is fed by the trace id chatCore
mints for each attempt, which also pairs the row with its `request.started`
dashboard event (diegosouzapw#13481). That id was `randomUUID().slice(0, 6)` — 6 hex
chars, 24 bits, 16,777,216 values.

The resulting collision is birthday-shaped against the rows already stored,
not a concurrency race between simultaneous inserts: with N rows in the table
every insert collides with probability N/2^24, which is ~1.6% at 268k rows and
matches the reported handful of failures per hour under ordinary traffic. The
losing insert throws inside `saveCallLogOperation`, which logs and swallows it,
so the row disappears with no signal to the caller — call analytics, combo
provider stats and cost rollups undercount with nothing but a container log
line to show for it.

Move the generator into `chatCore/traceId.ts` so the constraint is stated
where the value is produced, and widen it to 16 hex chars (64 bits). That
matches `randomHex(16)` in grok-web and `randomId(16)` in the OTEL exporter,
keeps log lines short, and takes the same 268k-row collision probability to
roughly 1.5e-14 per insert.

Tests reproduce the silent drop through the real save path (the second write
is lost and `saveCallLog` still resolves), assert the id width and charset,
and draw 200k ids without a duplicate. Reverting the width to 6 fails the
latter two with a real collision; dropping the dash-strip fails the charset
assertion.

Fixes diegosouzapw#14338
sxh313 added a commit to sxh313/OmniRoute that referenced this pull request Sep 21, 2026
PR diegosouzapw#14341, filed eleven minutes before mine for the same issue, adds
tests/unit/call-log-id-collision-14338.test.ts for the trace-id widening half of diegosouzapw#14338. That is a
different fix, not a duplicate, but the identical file path would conflict whichever of the two merges
second, so this one is renamed. No behaviour change.
@abhisheksharma2411

Copy link
Copy Markdown
Contributor Author

CI is red on this PR and none of it comes from this change. I chased it properly rather than asserting it, because the sibling PR #14356 (same base, dashboard files only) is fully green, which made "pre-existing" look like a weak excuse.

The diff touches four files: open-sse/handlers/chatCore.ts, the new chatCore/traceId.ts, one new test, and a changelog.d entry.

Why the two PRs differ. The fast path is impact-selected — Fast Quality Gates prints a "Candidate tests absent from this diff" section listing open-sse/handlers/chatCore.ts (traceId) -> … fan-out. A change to chatCore.ts pulls in a much wider test set than a dashboard-only change, so this PR runs tests #14356 never does. Those extra tests include the repo's current base-reds.

Verified by running the failures on a clean release/v3.8.51 worktree, with none of my changes:

test on clean base
i18n pt-BR integrity ✖ fails
should contain every key present in en.json (no drift, #6695) ✖ fails
provider detail client entry stays browser-bundle safe ✖ fails
every local import of every npm-shipped wrapper survives the prune (allowlist) ✖ fails
every local import … is enforced by check:pack-artifact ✖ fails
tracked source and documentation surfaces have no retired asset reference ✖ fails
handleChat uses the emergency fallback model on budget exhaustion ✖ fails
handleChat returns the primary budget error when emergency fallback also fails ✖ fails
handleUcVideoGeneration (persona) times out with 504 when never ready ✖ fails

Nine of eleven, failing with nothing of mine applied. The browser-bundle safe one is the EditConnectionModal → ClaudeConnectionFields missing-export breakage, which is independent of this PR.

The remaining two are not mine either:

  • #13459 saveRequestUsage stores the canonical provider id for an alias — passes in isolation on this branch (1 pass / 0 fail). It only fails inside the shard, and this PR adds a test file, which shifts --test-shard=2/4 membership and changes execution order. That is cross-test pollution surfaced by re-partitioning, not a behaviour change.
  • maxWaitMs=0 … — passes in isolation on the base; timing-sensitive.

Worth flagging separately: shard membership is a function of the file list, so adding any test file re-partitions all four shards and can surface order-dependent failures that have nothing to do with the change. That is a property of --test-shard worth knowing when reading a red matrix on a PR that adds tests.

Local verification of the change itself is unchanged from the description: the new suite is 3 pass / 0 fail, mutation-tested 3/3, and call-log-file-rotation's orphan-cleanup failure reproduces identically on a stashed tree.

Happy to rebase if #14331 (the base-red drain) lands first and you'd rather see a clean matrix.

@diegosouzapw

Copy link
Copy Markdown
Owner

Your diagnosis landed first and it was the correct one — the 24-bit id space, id TEXT PRIMARY KEY,
the birthday bound against stored rows rather than concurrent inserts, and the explicit "this is not
a race" — roughly 14 hours before the canonical issue (#14451) was filed with the same conclusion.
That deserves credit and will get it: whichever fix lands will carry a Co-authored-by trailer for
you.

On the code, I'm recommending #14474 as the one to merge: it decouples the primary key from the
trace id instead of widening it, and it fixes generateLogId() at the same time, which leaves
#13099's surface closed too. Two things about this branch that pushed the decision that way: the
trace id remains the primary key here, so the column keeps carrying two concerns, and the first test
asserts the silent row drop as expected behaviour rather than as the bug. All three PRs edit
chatCore.ts:549 and callLogs.ts:139-144, so only one can land as-is. Nothing here is wasted —
thank you for the analysis, it is what made the mechanism clear.

HouMinXi added a commit to HouMinXi/OmniRoute that referenced this pull request Sep 22, 2026
…attempts

When requests retry or fallback, a 6-character trace id prefix used as
call_logs.id collides with rows already stored, and an explicit primary
key that is already taken drops the new row.

- Use full randomUUID for request trace identifiers in chatCore.
- In saveCallLog, catch UNIQUE on an explicit id and retry with a
  generated UUID instead of dropping the row.
- persistAttemptLogs no longer keys the row on traceId. The dashboard
  correlation token is unchanged.

The diagnosis (24-bit id space, id TEXT PRIMARY KEY, birthday collision
against stored rows, not a concurrency race) was already in diegosouzapw#14341.

Fixes diegosouzapw#14451

Co-authored-by: Abhishek Sharma <abhicse24@gmail.com>
Signed-off-by: Minxi Hou <houminxi@gmail.com>
@abhisheksharma2411

Copy link
Copy Markdown
Contributor Author

Agreed on all of it — closing in favour of #14474, and thank you for the credit offer, though the analysis being useful is enough.

Your second point is the one worth me writing down, because it's a defect in how I wrote the test rather than a preference:

the first test asserts the silent row drop as expected behaviour rather than as the bug

That's correct, and it's worse than it looks. The test is called "a duplicate call-log id is dropped silently, not surfaced" and it asserts rows.length === 1, model === "model-a", status === 200 — the second row's loss, pinned as a contract. A characterization test written in the grammar of a specification. The consequence is that #14474 makes it fail, and it fails looking like a regression when it is in fact the fix landing. If that test had reached main ahead of the real repair it would have been an argument against repairing it.

The lesson I'm taking: a test that documents current-broken behaviour needs to say so in its name and its assertion message, or it becomes a defence of the bug. Something like assert.equal(rows.length, 1, "CHARACTERIZING #14338 — this is the defect; delete this test when the id is decoupled") would have made it self-cancelling.

Your first point stands too. Widening 24 → 64 bits takes the collision probability from ~1.6% at 268k rows to ~1.5e-14, which fixes the arithmetic while leaving call_logs.id doing two jobs — a row key and a live-topology correlation token — and those two have genuinely different requirements. #14474 separates them, which is why it also closes #13099 and mine doesn't.

I've left a review on #14474. One finding there worth flagging here since it touches the same seam: correlationId || traceId means the traceId → row link survives only when the caller passed no correlationId, and every test in that PR passes null. Probed with a control; details are on that PR.

Closing this. No wasted effort from my side — thanks for reading the analysis carefully enough to find the test problem in it.

HouMinXi added a commit to HouMinXi/OmniRoute that referenced this pull request Sep 24, 2026
…attempts

When requests retry or fallback, a 6-character trace id prefix used as
call_logs.id collides with rows already stored, and an explicit primary
key that is already taken drops the new row.

- Use full randomUUID for request trace identifiers in chatCore.
- In saveCallLog, catch UNIQUE on an explicit id and retry with a
  generated UUID instead of dropping the row.
- persistAttemptLogs no longer keys the row on traceId. The dashboard
  correlation token is unchanged.

The diagnosis (24-bit id space, id TEXT PRIMARY KEY, birthday collision
against stored rows, not a concurrency race) was already in diegosouzapw#14341.

Fixes diegosouzapw#14451

Co-authored-by: Abhishek Sharma <abhicse24@gmail.com>
Signed-off-by: Minxi Hou <houminxi@gmail.com>
HouMinXi added a commit to HouMinXi/OmniRoute that referenced this pull request Sep 24, 2026
…attempts

When requests retry or fallback, a 6-character trace id prefix used as
call_logs.id collides with rows already stored, and an explicit primary
key that is already taken drops the new row.

- Use full randomUUID for request trace identifiers in chatCore.
- In saveCallLog, catch UNIQUE on an explicit id and retry with a
  generated UUID instead of dropping the row.
- persistAttemptLogs no longer keys the row on traceId. The dashboard
  correlation token is unchanged.

The diagnosis (24-bit id space, id TEXT PRIMARY KEY, birthday collision
against stored rows, not a concurrency race) was already in diegosouzapw#14341.

Fixes diegosouzapw#14451

Co-authored-by: Abhishek Sharma <abhicse24@gmail.com>
Signed-off-by: Minxi Hou <houminxi@gmail.com>
HouMinXi added a commit to HouMinXi/OmniRoute that referenced this pull request Sep 24, 2026
…attempts

When requests retry or fallback, a 6-character trace id prefix used as
call_logs.id collides with rows already stored, and an explicit primary
key that is already taken drops the new row.

- Use full randomUUID for request trace identifiers in chatCore.
- In saveCallLog, catch UNIQUE on an explicit id and retry with a
  generated UUID instead of dropping the row.
- persistAttemptLogs no longer keys the row on traceId. The dashboard
  correlation token is unchanged.

The diagnosis (24-bit id space, id TEXT PRIMARY KEY, birthday collision
against stored rows, not a concurrency race) was already in diegosouzapw#14341.

Fixes diegosouzapw#14451

Co-authored-by: Abhishek Sharma <abhicse24@gmail.com>
Signed-off-by: Minxi Hou <houminxi@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(backend): call_logs.id UNIQUE constraint race on concurrent inserts

2 participants