Skip to content

fix(dashboard): retain token usage for in-memory request rows - #14324

Merged
diegosouzapw merged 2 commits into
diegosouzapw:release/v3.8.51from
woodsonl:fix/in-memory-token-usage
Sep 24, 2026
Merged

diegosouzapw merged 2 commits into
diegosouzapw:release/v3.8.51from
woodsonl:fix/in-memory-token-usage

Conversation

@woodsonl

@woodsonl woodsonl commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Fixes completed requests showing 0 / 0 tokens in the Requests dashboard when provider usage is available but the dashboard is displaying an in-memory row.

Reproduction

On September 21, three successful OpenCode requests through an OpenAI-compatible endpoint showed zero tokens in the dashboard:

Request start (UTC) Dashboard input / output Provider input / output
14:43:35 0 / 0 43,979 / 243
14:43:43 0 / 0 48,935 / 250
14:43:50 0 / 0 50,613 / 404

The API rows had status: 200, active: false, completed: true, and detailState: "in-memory". The streams completed and the client continued sending requests.

The deployment had request-history storage disabled in a custom maintenance image. Its OMNIROUTE_SAVE_REQUEST_HISTORY flag is specific to that image. Upstream can show the same incorrect counts whenever it serves an in-memory row without a matching persisted row, including while a write is pending.

Cause and changes

persistAttemptLogs() passes provider usage to saveCallLog(), but the pending and completed request caches did not receive those counts. buildCallLogListRows() hardcoded tokens: { in: 0, out: 0 } for both kinds of in-memory row.

This change copies the counts into the existing request cache before checking whether disk writes are allowed. The dashboard then reads them from the cached detail.

persistAttemptLogs()
  → saveCallLog()
      → normalize token counts
      → updateRequestTokensById(pendingRequestId ?? id)
      → existing disk-write path
  → buildCallLogListRows() reads detail.tokens

The IDs need separate handling. Persisted attempts use traceId; live requests use pendingRequestId. These can differ during combo routing. attemptLogging.ts now passes both, so the cache update reaches the right request without changing the persisted attempt ID.

File Change
open-sse/handlers/chatCore/attemptLogging.ts Pass pendingRequestId alongside the log's traceId.
src/lib/usage/callLogs.ts Normalize counts with the existing tokenAccounting helpers and update the live detail before the disk-write check.
src/lib/usage/usageHistory.ts Store typed token counts on pending details. Update an existing completed detail if usage arrives after completion. Unknown IDs create no entries.
src/app/api/usage/call-logs/route.ts Return detail.tokens when available.
src/lib/usage/completedRequestDetails.ts Keep newer token counts when background artifact loading finishes.
tests/unit/call-log-in-memory-usage.test.ts Cover the logging, cache, and dashboard paths together.

The cached fields are in, out, cacheRead, cacheCreation, reasoning, and compressed. The existing helpers handle Anthropic input/cache totals. Explicit zero output stays zero; missing optional cache/reasoning counts stay null.

Review of the first revision found a race: background artifact loading could write an older detail back to the cache and erase usage received after completion. Commit 2052a5e7 preserves the current token counts when saving the loaded detail. A regression test failed with undefined !== 4321 before that change.

Tests

The tests use an isolated data directory and real pending-request state. They close the call-log write gate, send usage through saveCallLog() or persistAttemptLogs(), and read the result through the dashboard row builder. The no-persistence cases assert that no call_logs row was written. No provider credentials or custom deployment flags are needed.

The first two tests failed with 0 !== 43979 before the fix. The five new cases cover:

  1. Usage received before completion.
  2. Usage received after completion.
  3. Concurrent requests sharing a model/provider/connection, missing usage updates, zero output, and Anthropic cache totals.
  4. Chat logging with different traceId and pendingRequestId values.
  5. Background artifact loading finishing after a token update.
node --import tsx/esm --test \
  tests/unit/call-log-in-memory-usage.test.ts \
  tests/unit/chatcore-attempt-logging.test.ts \
  tests/unit/call-log-provider-display.test.ts \
  tests/unit/call-log-detailed-tokens.test.ts \
  tests/unit/call-logs-row-filter.test.ts \
  tests/unit/call-log-save-drain.test.ts \
  tests/unit/completed-detail-pipeline-precedence.test.ts
npm run typecheck:core
npm run check:cycles

Locally, all 39 focused tests and the core typecheck passed after the race fix. The cycle check passed on the first revision; the follow-up adds no imports. Changed-file ESLint, Prettier, and pre-commit checks passed. Full local suites and coverage were not run.

All five new tests also passed in CI for 2052a5e7, along with Vitest fast-path and both Semgrep checks.

CI failures

CI is still red. The first run reported the failures below; the follow-up run repeats the API errors and has failures in the other listed check groups.

Check Failure from inspected logs
API Route Typecheck TS2677 in src/app/api/v1/_shared/rerankProviderNodes.ts and src/mitm/handlers/antigravity.ts.
Merge integrity Generated omni-version-manager skill is out of sync.
Docs Gates Missing .env.example entries for BRIDGE_PORT, CERT_DIR, NEXT_PUBLIC_PORT, OPENWA_SERVICE_PORT, and ROUTER_URL.
ESLint Errors in cliRuntime.ts, arena/compression tests, and translator-openai-to-gemini.test.ts.
Fast Quality Gates Mutation test inventory, pack policy, and dashboard typecheck errors in proxySaveRefresh.test.tsx, antigravity.ts, and Sidebar.tsx.
Unit tests Failures in provider catalogs, CLI discovery, packaging, environment contracts, and credential inventory. The five tests added here passed.

The named API files and the checked lint files are unchanged from the base. #13866 tracks release-branch failures. Each failed test has not been independently reproduced on the base. The local API gate also failed with the two TS2677 errors.

⚠️ base-red inherited: #13866

Limits and deployment check

The change stores token counts in the existing completed-request cache, which expires entries after 120 seconds and is capped at 256 entries and 16 MiB of estimated payload size. It adds no request/response payload storage or schema changes. Existing filtering and persisted-row precedence remain in place.

Rows still show the existing 0 / 0 fallback when usage has not arrived. This PR does not add an “unavailable” display or estimate usage from response text.

After deployment, compare a successful detailState: "in-memory" row from /api/usage/call-logs with the same request's upstream usage. Check both a short and a streamed completion before the two-minute cache expiry. Input/output and reported cache/reasoning counts should agree. The patch has not been deployed or tested this way yet.

Related: #14204 / #14205 fix usage parsing in the OpenAI-to-Claude translator. This PR handles counts that were received but omitted from the live dashboard rows.

@woodsonl

woodsonl commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor Author

Pushed 2052a5e7 to fix a race found during review: background artifact loading could overwrite token counts received after request completion. The new regression test failed before the fix and now passes.

All five new tests passed in CI. The local focused run passed 39 tests. Vitest fast-path and both Semgrep checks also passed.

CI still has failures. The API check repeats TS2677 in unchanged rerankProviderNodes.ts and antigravity.ts; the unit shard containing these tests fails elsewhere. The PR description lists the failures and links to the logs. #13866 tracks the release-branch problems, but not every failure has been reproduced on the base.

@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks for the detailed repro and the write-up connecting persistAttemptLogs() →
saveCallLog() → the dashboard's buildCallLogListRows(). Ran your new
call-log-in-memory-usage.test.ts at HEAD — all 5 cases pass. The fix is scoped exactly to
the files you listed and doesn't touch the disk-write path. Looks ready to merge; the
failing CI checks on this PR match the release-tip's currently-inherited base-red
(#14547/#14496), not anything in your diff.

@diegosouzapw
diegosouzapw merged commit 95651e6 into diegosouzapw:release/v3.8.51 Sep 24, 2026
7 of 16 checks passed
diegosouzapw added a commit to HouMinXi/OmniRoute that referenced this pull request Sep 25, 2026
…request

The primary-key change dropped pendingRequestId from the saveCallLog entry.
saveCallLog uses it (not the row id) to update token usage on the live
in-memory request row (diegosouzapw#14324), so the dashboard lost the counters of every
chat attempt. call-log-in-memory-usage.test.ts 'chat attempt logging connects
its trace id to the live pending request id' fails without this line.
thinh0704hcm added a commit to thinh0704hcm/OmniRoute that referenced this pull request Sep 25, 2026
… fix

Upstream sync (213 commits) plus fix(responses): synthesize SSE directly from native Responses JSON on forced non-stream. Conflict resolutions: docs/env/pre-commit take upstream; attemptLogging takes upstream UUID-PK (diegosouzapw#14324); upstreamTimeouts keeps upstream abort-link lifetime; virtualFactory takes upstream rawCustomModels filter; runtimeUnits combines first-content deadline race with upstream signal threading; rateLimitManager takes upstream queue/execution rewrite (covers fork cancellation intents); stream.ts and search route combine additive fields/blocks; autoRouting combines invocation tracing with builtin hardening; redaction fixture and attempt-logging tests take upstream expectations.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants