fix(dashboard): retain token usage for in-memory request rows - #14324
diegosouzapw merged 2 commits into
Conversation
|
Pushed All five new tests passed in CI. The local focused run passed 39 tests. Vitest fast-path and both Semgrep checks also passed. CI still has failures. The API check repeats TS2677 in unchanged |
|
Thanks for the detailed repro and the write-up connecting |
95651e6
into
diegosouzapw:release/v3.8.51
…request The primary-key change dropped pendingRequestId from the saveCallLog entry. saveCallLog uses it (not the row id) to update token usage on the live in-memory request row (diegosouzapw#14324), so the dashboard lost the counters of every chat attempt. call-log-in-memory-usage.test.ts 'chat attempt logging connects its trace id to the live pending request id' fails without this line.
… fix Upstream sync (213 commits) plus fix(responses): synthesize SSE directly from native Responses JSON on forced non-stream. Conflict resolutions: docs/env/pre-commit take upstream; attemptLogging takes upstream UUID-PK (diegosouzapw#14324); upstreamTimeouts keeps upstream abort-link lifetime; virtualFactory takes upstream rawCustomModels filter; runtimeUnits combines first-content deadline race with upstream signal threading; rateLimitManager takes upstream queue/execution rewrite (covers fork cancellation intents); stream.ts and search route combine additive fields/blocks; autoRouting combines invocation tracing with builtin hardening; redaction fixture and attempt-logging tests take upstream expectations.
Fixes completed requests showing
0 / 0tokens in the Requests dashboard when provider usage is available but the dashboard is displaying an in-memory row.Reproduction
On September 21, three successful OpenCode requests through an OpenAI-compatible endpoint showed zero tokens in the dashboard:
The API rows had
status: 200,active: false,completed: true, anddetailState: "in-memory". The streams completed and the client continued sending requests.The deployment had request-history storage disabled in a custom maintenance image. Its
OMNIROUTE_SAVE_REQUEST_HISTORYflag is specific to that image. Upstream can show the same incorrect counts whenever it serves an in-memory row without a matching persisted row, including while a write is pending.Cause and changes
persistAttemptLogs()passes provider usage tosaveCallLog(), but the pending and completed request caches did not receive those counts.buildCallLogListRows()hardcodedtokens: { in: 0, out: 0 }for both kinds of in-memory row.This change copies the counts into the existing request cache before checking whether disk writes are allowed. The dashboard then reads them from the cached detail.
The IDs need separate handling. Persisted attempts use
traceId; live requests usependingRequestId. These can differ during combo routing.attemptLogging.tsnow passes both, so the cache update reaches the right request without changing the persisted attempt ID.open-sse/handlers/chatCore/attemptLogging.tspendingRequestIdalongside the log'straceId.src/lib/usage/callLogs.tstokenAccountinghelpers and update the live detail before the disk-write check.src/lib/usage/usageHistory.tssrc/app/api/usage/call-logs/route.tsdetail.tokenswhen available.src/lib/usage/completedRequestDetails.tstests/unit/call-log-in-memory-usage.test.tsThe cached fields are
in,out,cacheRead,cacheCreation,reasoning, andcompressed. The existing helpers handle Anthropic input/cache totals. Explicit zero output stays zero; missing optional cache/reasoning counts stay null.Review of the first revision found a race: background artifact loading could write an older detail back to the cache and erase usage received after completion. Commit
2052a5e7preserves the current token counts when saving the loaded detail. A regression test failed withundefined !== 4321before that change.Tests
The tests use an isolated data directory and real pending-request state. They close the call-log write gate, send usage through
saveCallLog()orpersistAttemptLogs(), and read the result through the dashboard row builder. The no-persistence cases assert that nocall_logsrow was written. No provider credentials or custom deployment flags are needed.The first two tests failed with
0 !== 43979before the fix. The five new cases cover:traceIdandpendingRequestIdvalues.Locally, all 39 focused tests and the core typecheck passed after the race fix. The cycle check passed on the first revision; the follow-up adds no imports. Changed-file ESLint, Prettier, and pre-commit checks passed. Full local suites and coverage were not run.
All five new tests also passed in CI for
2052a5e7, along with Vitest fast-path and both Semgrep checks.CI failures
CI is still red. The first run reported the failures below; the follow-up run repeats the API errors and has failures in the other listed check groups.
src/app/api/v1/_shared/rerankProviderNodes.tsandsrc/mitm/handlers/antigravity.ts.omni-version-managerskill is out of sync..env.exampleentries forBRIDGE_PORT,CERT_DIR,NEXT_PUBLIC_PORT,OPENWA_SERVICE_PORT, andROUTER_URL.cliRuntime.ts, arena/compression tests, andtranslator-openai-to-gemini.test.ts.proxySaveRefresh.test.tsx,antigravity.ts, andSidebar.tsx.The named API files and the checked lint files are unchanged from the base. #13866 tracks release-branch failures. Each failed test has not been independently reproduced on the base. The local API gate also failed with the two TS2677 errors.
Limits and deployment check
The change stores token counts in the existing completed-request cache, which expires entries after 120 seconds and is capped at 256 entries and 16 MiB of estimated payload size. It adds no request/response payload storage or schema changes. Existing filtering and persisted-row precedence remain in place.
Rows still show the existing
0 / 0fallback when usage has not arrived. This PR does not add an “unavailable” display or estimate usage from response text.After deployment, compare a successful
detailState: "in-memory"row from/api/usage/call-logswith the same request's upstream usage. Check both a short and a streamed completion before the two-minute cache expiry. Input/output and reported cache/reasoning counts should agree. The patch has not been deployed or tested this way yet.Related: #14204 / #14205 fix usage parsing in the OpenAI-to-Claude translator. This PR handles counts that were received but omitted from the live dashboard rows.