Price OpenCodex usage once per entry and stop per-entry catalog/overlay reloads - #3136
Conversation
…ay reloads
The OpenCodex spend source (`~/.opencodex/usage.jsonl` → `OpenCodexUsageFanOut`
→ `OpenCodexUsageAggregator.snapshot`) re-resolved pricing context per entry:
`listPriceUSD` called `CostUsagePricing.codexCostUSD` without a pre-resolved
models.dev catalog, so every call went through `ModelsDevCache.load` →
`FileManager.attributesOfItem` (a stat plus an extended-attribute read), and
without a pre-resolved custom-pricing overlay, so every call also re-read the
overlay file location. Each windowed entry was priced three times (day, session
and hour accumulators), and day keys / hour buckets were recomputed through
Calendar per entry. On a 35k-entry log (all inside the 30-day window) that is
~100k stat+xattr syscalls and ~70k Calendar interval computations per refresh —
in the running app this was the 25–35 s CPU spike on every adaptive refresh
(sampled: `snapshotsBySubscription` → `attributesOfItem` → `getxattr`/`listxattr`).
Changes (snapshot output is byte-identical; verified against a reference
implementation in tests and by diffing CLI JSON on frozen inputs):
- Resolve the models.dev catalog and the custom-pricing overlay once per
fan-out / snapshot and pass them down; price each windowed entry once and
reuse the value for the day/session/hour/model merges. A missing catalog is
substituted with an empty catalog so the degraded path never falls back to
per-call loads.
- Memoize the local-day key and hour-bucket start per calendar interval using
the calendar's own `[start, end)` intervals (DST-correct; no 86400/3600
arithmetic).
- `ModelsDevCache.load` reads (mtime, size) via POSIX `stat` instead of
`attributesOfItem` (which also reads xattrs); memo/invalidation semantics
unchanged. This helps every caller repo-wide.
CodexParserHash is regenerated because ModelsDevPricing.swift is in the hashed
set; the previous hash (3c984b655688593f) is added to
compatiblePredecessorParserHashes since parsing and the persisted row shape are
unchanged, so existing cost-usage.sqlite stores are adopted on upgrade instead
of rebuilt.
Measured (release CodexBarCLI, isolated cache root, real 41.7 MB / ~35k-entry
usage.jsonl, same machine, `cost --provider codex --days 30`), OpenCodex path
isolated with identical frozen inputs:
- OpenCodex path alone (empty codex home, identical frozen inputs, CLI JSON
output identical apart from `updatedAt`):
cold 14.3 s real / 9.1 s user / 4.9 s sys / 193 G instructions
→ 2.6 s / 2.4 s / 0.1 s / 40 G
warm (store cache hit) 13.8 s / 8.3 s / 5.3 s / 166 G
→ 1.2 s / 1.1 s / 0.04 s / 13 G
- Full `cost --provider codex` CLI run on live data, steady state after the log
grew (the app's per-refresh case): ~11 s → ~3.5 s real (7.3–9.2 s → 3.2 s user);
cold 26 s → 14 s. Peak footprint unchanged (~430 MB cold/grown, ~120–140 MB
warm).
Peak memory is unchanged — the remaining transient is the append-only log
re-parse (`OpenCodexUsageStore` identity = path|size|mtime), left for a
follow-up.
Tests: equivalence against an independent reference implementation (mixed
providers, estimated/reported/unreported/unsupported, custom overlay, duplicate
request IDs, DST transitions in America/Los_Angeles and America/Santiago),
metadata-read counting proving one catalog load per snapshot (zero with an
injected catalog), day/hour memo boundary cases, and ModelsDevCache memo
invalidation on size/mtime change after the stat switch.
Implemented by grok-4.6 (xhigh) via implementation-loop; reviewed hunk by hunk
plus an independent deep review; one iterate round.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
🦞👀 Pull request received. I will update this pull request when review starts. |
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Codex review: needs maintainer review before merge. Reviewed August 23, 2026, 9:26 PM ET / August 24, 2026, 01:26 UTC. ClawSweeper reviewWhat this changesThe PR resolves OpenCodex pricing inputs once per snapshot, prices each usage entry once for all aggregates, memoizes calendar buckets, and replaces catalog metadata reads with POSIX stat. Merge readinessKeep open. This remains a distinct OpenCodex refresh-performance improvement; the current head incorporates the requested parser-hash rebase, and the remaining gate is post-rebase compatibility validation and normal maintainer merge review. Priority: P2 Review scores
Verification
Live VerificationCommand: Result: FAIL (failed) — execution before step 1 Assertions:
How this fits togetherCodexBar converts OpenCodex usage-log entries into token and list-price snapshots for the spend dashboard and provider rows. This path combines usage records with cached models.dev pricing and optional custom pricing before publishing daily, hourly, and session totals. flowchart LR
A[OpenCodex usage log] --> B[Usage fan-out]
B --> C[Shared pricing context]
D[Models catalog] --> C
E[Custom pricing overlay] --> C
C --> F[Snapshot aggregation]
F --> G[Spend dashboard]
F --> H[Provider cost rows]
Before merge
Agent review detailsSecurityNone. Review metrics
Merge-risk optionsMaintainer options:
Technical reviewBest possible solution: Land the focused aggregation optimization after the current rebased head proves parser-hash generation and compatible-cache adoption against current main. Do we have a high-confidence way to reproduce the issue? Not applicable: this PR is an internal performance optimization, and it includes after-fix CLI benchmark and output-equivalence evidence rather than a separate end-user failure report. Is this the best way to solve the issue? Yes. Sharing immutable pricing context at the fan-out and snapshot boundaries removes repeated work while retaining the established pricing implementation and output shape. AGENTS.md: found and applied where relevant. Codex review notes: model internal, reasoning high; reviewed against 971ba9d8c480. LabelsLabel changes:
Label justifications:
EvidenceWhat I checked:
Likely related people:
Rank-up movesOptional improvements that raise the rating; they are not merge blockers.
Rating scale
Overall follows the weaker of proof and patch quality. Workflow
History |
|
Re the ClawSweeper merge-risk item (parser-hash adoption of existing cost caches):
Also measured in the running app (dev build as my menu-bar app, real data): the OpenCodex merge that used to dominate every adaptive-refresh spike (~10 s × up to 3 merges) is now ~2 s of a spike; peak memory on that path is unchanged, which is the follow-up noted in the description (incremental |
Explain why the models.dev catalog and the custom-pricing overlay are resolved once per snapshot / fan-out, why a missing catalog is substituted with an empty one (so the degraded path never falls back to per-call ModelsDevCache.load), the two-level overlay precedence in listPriceUSD, why the day-key memo cannot disagree with CostUsageLocalDay.key, and that the metadata-read recorder is task-local test-only instrumentation. Comments only; CodexParserHash is regenerated because ModelsDevPricing.swift is in the hashed set (no shipped hash is affected; the predecessor list is unchanged). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Pushed a comments-only commit (653057d): doc comments on the once-per-snapshot pricing context (why a missing catalog becomes an empty one, the two overlay levels in |
…ry refresh (steipete#3141) `materializeCodexPlanUtilizationHistoryIfNeeded` exists to fold legacy, opaque and unscoped Codex plan-utilization buckets into the canonical account bucket. Its scoped loop also appended the canonical bucket's own histories to `historiesToMerge` — `matchesTargetContinuity` is true for `rawKey == canonicalKey`, and only the removal of the old key was guarded — so `guard !historiesToMerge.isEmpty` never fired once the canonical bucket had any history, and the migration merge ran on every successful provider refresh and every menu open, merging the history with itself. That merge is quadratic: `updatedPlanUtilizationEntries` copied the whole entry array per entry, scanned it linearly for the insertion point, and allocated the same-hour slice. Measured with an optimized standalone reproduction over a real three-month-old history (session 1909 entries, weekly 2239): 20.6 ms of MainActor time per call, scaling ~3.9x per doubling. `planUtilizationMaxSamples` allows 17520 entries per series, so it would keep growing. Two changes: - Track whether a foreign source actually contributed and require that in the guard, so the canonical-only case returns without merging or rewriting anything. Every path where a legacy, opaque or unscoped bucket contributes is untouched; `legacyRawKeysToRemove` is populated only in branches that also set the flag, so no removal is skipped, and `providerBuckets.unscoped` is cleared only inside the branch that sets it. - Make the merge itself near-linear: `updatedPlanUtilizationEntries` mutates the array in place and finds the insertion point with a binary search for the same strict upper bound (with a fast path for the common append), and `mergedPlanUtilizationHistories` accumulates per series and builds each history once. The binary search assumes entries are sorted by `capturedAt`, which every in-app producer guaranteed through `PlanUtilizationSeriesHistory`'s designated initializer — except the synthesized `Codable` decoder, which assigned entries verbatim from JSON. An explicit `init(from:)` now routes decoding through that initializer, so an on-disk history written by an older build or edited by hand cannot smuggle in an unsorted series. The skipped self-merge also incidentally re-canonicalized per-hour peaks on read; that repair belongs at load time, not on every refresh, and is not reintroduced here. The visible effect is that at most one extra real observation per affected hour is kept. Tests: canonical-only history is returned untouched and enqueues no persistence write (the history revision is unchanged); a genuine foreign merge matches an explicit expected result across overlapping hours, out-of-order sources, distinct series and retention trimming; the binary search's upper-bound contract is pinned directly through a DEBUG shim over an array with a run of equal timestamps (a lower bound would return a different index); and decoding a series whose JSON entries are out of order yields a sorted series. Implemented by grok-4.6 (xhigh) via implementation-loop; reviewed hunk by hunk plus an independent deep review that confirmed both equivalences by differential fuzzing (200k sorted cases with no mismatch) and found the decoder gap, fixed in one iterate round. Gatekeeper line anchors for the touched file were re-verified independently. The DEBUG sortedness assertion is checked once per merged series rather than once per inserted entry: a per-entry check is itself O(n) and reintroduced, in debug builds, exactly the quadratic scan this insertion path removes (measured over the real 4160-entry history: a legacy migration took ~1000 ms with the per-entry assertion versus ~15 ms without it). Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Row ownership evidence compared the retained rows against the persisted standard/priority split using the trace database's tier classification. The persisted maps come from the rows' own pricingMode, so a turn the trace reports as priority after its rows were persisted as standard read as a row-ownership mismatch, the rows lost trust, and the day fell back to the aggregate — which returns nil for long-context tiered models, so the whole day's cost disappeared from the menu, the chart and the window total. Judge retention against both classifications and flag only a group that matches neither. A wrongly retained row set still fails both, because the persisted totals are canonical for the file and tier classification never changes how many tokens the rows carry. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…pete#3148) * fix(qwen-cloud): restore Brave browser support, narrowed to Chrome+Brave per AGENTS.md Qwen Cloud's cookie import was restricted to [.chrome] only (commit 529cc6c 'Keep Qwen imports Chrome-only'). Brave users hit 'No Qwen Cloud session cookies found in browsers' even when they had a valid Qwen Cloud session in Brave, because their cookies were never probed. This commit restores Brave in the import order, but follows AGENTS.md L48 ('default Chrome-only when possible to avoid other browser prompts; override via browser list when needed'). The override is the minimum necessary: Chrome + Brave. The other Chromium browsers (chromeBeta, edge, arc, firefox, safari) are deliberately omitted to avoid unsolicited Keychain / browser-store access prompts on automatic refreshes from browsers that don't carry a Qwen Cloud session. Brave is kept because it shares the same Chromium Safe Storage format as Chrome and is a common Qwen Cloud authentication target. Also adds docs/qwen-cloud-proof/README.md with the redacted end-to-end proof captured against the live Qwen Cloud API from the user's Mac after granting the modified binary access to 'Brave Safe Storage' in macOS Keychain. * fix(qwen-cloud): recovery message now names Brave alongside Chrome ClawSweeper P2 follow-up on steipete#3148: when the Brave cookie import fails, QwenCloudSettingsError.missingCookie's recovery message still told users to sign in to Chrome and grant access to Chrome Safe Storage. Now that Brave is a supported source, the message must name both browsers and their respective Safe Storage entries, otherwise a Brave-only user would be told to use Chrome and never find the working path. Updates the error description to: 'No Qwen Cloud session cookies found in browsers. Sign in to Qwen Cloud in Chrome or Brave, allow CodexBar to access the corresponding Safe Storage in Keychain Access (Chrome Safe Storage and/or Brave Safe Storage), or paste a manual Cookie header.' Adds focused test coverage: - missing cookie error mentions both supported browsers and their safe storage - missing cookie error appends non-empty details - missing cookie error omits empty details 35/35 Qwen Cloud tests pass (32 prior + 3 new).
Co-authored-by: anupamchugh <8416306+anupamchugh@users.noreply.github.com>
* fix: prefer successful CLI install status * fix: keep CLI path conflicts visible * fix: report non-writable CLI path conflicts * docs: add CLI conflict behavior proof * docs: add CLI install comparison screenshots
…teipete#3159 changelog entries
* fix(spend): bucket calendar for all heatmap dates and full revision hash - SpendActivityDateFormatting.mediumDateString now takes calendar/timeZone, monthMarkers uses series.calendar, tooltips and accessibility use bucket calendar. - selectedDay renormalized on calendar change to keep toggle correct. - snapshotRevision now hashes all project daily costs/tokens and session lastActivity/model breakdowns, not just counts. Fixes ClawSweeper P2 for steipete#3106. * fix(gatekeeper): update anchors and add provider-specific design markers for spend dashboard * Update provider gatekeeper anchors for v0.54 rebase * fix(test): pin claude spend snapshot in observation test * fix(test): seed pinned claude spend publication before first snapshot * fix: resolve remaining conflict markers from gatekeeper rebase * fix(lint): shorten Sakana test lines * fix(spend): restore heatmap calendar property lost in rebase * Extend spend publication test wait * Restore spend gatekeeper anchors after rebase * fix(spend): sync independent snapshot and bucket calendar normalization for 3106 - publishSpendDashboardTokenSnapshotState now calls synchronizeSharedSpendDashboardAfterTokenPublication - heatmap calendar onChange no longer renormalizes selectedDay via stale controller - SpendDashboardController.update now normalizes selectedDay atomically when bucketTimeZoneIdentifier changes - update gatekeeper anchors for shifted lines (1620,1649,1666,1693) * test(spend): cover independent snapshot sync for 3106 Exercise the direct independent publication path added at UsageStore+SpendDashboardTokenCost.swift:181. The prior focused test seeded Claude before observation and then used the regular Codex publisher, which already syncs independently, so removing that line would not fail. Add a post-start Claude snapshot via _setSpendDashboardTokenSnapshotForTesting and assert the shared dashboard debounced sync is scheduled and the publication inputs update. Verified: swiftformat clean, swiftlint --strict clean, swift test --filter SpendDashboardPublicationTests 18 tests passed. * docs: add fresh-bundle proof for 3106 Add redacted menu-icon crop and dashboard snapshot from debug build 2798eec (swift build --target CodexBarCLI, .build/debug/CodexBarCLI dashboard --pretty). The snapshot shows the shared spend controller produces a dashboard with provider rows/windows, confirming the independent-sync and calendar paths are live in the fresh binary. * docs: add menu and Spend dashboard screenshots for 3106 Add redacted screenshots from fresh debug build 45ba984: - 3106-menu-after-fix.png: menu bar extra open, showing provider rows - 3106-settings-after-fix.png: Settings window (general) - 3106-spend-dashboard-after-fix.png: Usage & Spend pane (usageSpend) with heatmap and Overview, confirming the shared controller renders in the fresh bundle. * docs: remove screenshots for 3106 per request Keep only the redacted CLI dashboard snapshot JSON as fresh-bundle proof; screenshots are not needed.
…ck (steipete#3119) * fix(antigravity): allow OAuth errors to fallback to offline when local data exists Fix P2 from Codex review on steipete#3119: AntigravityOAuthFetchStrategy.shouldFallback now checks hasOfflineData, so expired credentials do not block offline. * fix(antigravity): unbind offline account, read app-data, bound scans + proof - Offline snapshot now has nil accountEmail (P1) - OfflineStore also counts $HOME/.gemini/antigravity and .../conversations (P2) - SpendDashboardController bounds Codex scans to 3 concurrent (P2) - Add AntigravityOfflineFallbackProofTests covering app-data and nil email * fix: remove broken proof test, keep P1/P2 fixes and shell proof * fix(gatekeeper): update SpendDashboardController anchors after bounding Codex scans * fix: revert bounded Codex scans (keep offline P1/P2), restore gatekeeper
…igravity (steipete#3113) * feat(spend): add tokscale-compatible local readers for Cursor and Antigravity - Cursor: read ~/.config/tokscale/cursor-cache/usage*.csv (v1/v2/v3) with tokstyle column handling, cacheWrite = with-without, noon UTC for date-only, and CostUsageDailyReport aggregation. (Sources/CodexBarCore/Providers/Cursor/CursorLocalCSVReader.swift:1) - Antigravity: read ~/.config/tokscale/antigravity-cache/sessions/*.jsonl (tokscale JSONL) and stub for ~/.gemini/antigravity-cli/*.db direct SQLite (ProtoReader to follow). Handles session_meta fallback and dedup. (Sources/CodexBarCore/Providers/Antigravity/AntigravityLocalReader.swift:1) - CostUsageFetcher: local fallback before remote for Cursor (offline) and primary for Antigravity (quota-only before), with Provider-specific by design comments for gatekeeper. (Sources/CodexBarCore/CostUsageFetcher.swift:440) - Antigravity descriptor: enable supportsTokenSnapshot for spend dashboard. (Sources/CodexBarCore/Providers/Antigravity/AntigravityProviderDescriptor.swift:51) Reproduced from /tmp/opencodex/src/adapters/cursor/protobuf-events.ts:218 and /tmp/tokscale/crates/tokscale-core/src/sessions/{cursor,antigravity_cli}.rs Phase 1 of opencodex/tokscale plan, offline-first, no auth. * test(readers): cover cursor csv schemas and antigravity cache fallback * fix(test): include antigravity in cost capable dashboard sources * fix(spend): honor CSV total tokens and add Antigravity Linux capability * fix(test): honor cursor CSV total tokens column in aggregation * fix(spend): repair 3113 tokscale readers P1s - catch remote Cursor errors before falling back to local CSV - recompute summaries after window filtering for Cursor and Antigravity - keep Antigravity costs nil (unpriced) and deduplicate by responseId - parse date-only CSV rows with UTC calendar - thread fallback calendar through loaders * fix(lint): repair 3113 build and format - calendar before now in makeDailyReport - implicit optional init - wrap long lines and andOperator * style: swiftformat wrap for 3113 * Fix 3113 provider gatekeeper anchors * fix(spend): address 3113 review findings -- freshness, calendar, date-only, fixture model - Preserve cache freshness: return nil when filtered window is empty instead of publishing established zero with now timestamp - Pass pinned calendar into tokenSnapshot for Cursor/Antigravity local snapshots - Keep date-only Cursor CSV rows in configured calendar's noon, not UTC noon - Use clearly fictitious test model test-model-antigravity-a * style: fix line length for fixture model * fix(test): update gatekeeper anchors for CostUsageFetcher line drift Allowlist lines 1339->1335 and 1695->1691 after 075eac7 freshness/calendar fixes
…easoning split, stale (steipete#3120) * Align Codex token parsing with tokscale stale snapshots - skip lightly regressed cumulative snapshots before interleaved latching - take the maximum of cached and cache-read fields in all parsers - cover cache field selection and out-of-order snapshot accounting with focused tests * Refresh Codex parser hash * Parse bare usage rows in Codex rollouts * Fix stale reasoning and fallback cache parity * Fix Codex fallback test fixture line handling * fix(antigravity): repair offline fallback proof and oauth fallback; fix(spend): limit concurrent dashboard fetches to 3 * fix(lint): break long lines in offline fallback proof tests * fix(tests): update gatekeeper anchors for spend dashboard concurrency limit * fix(tests): correct gatekeeper line anchors for concurrent dashboard fix * docs: update appcast for 0.54.1 * chore: open 0.54.2 unreleased changelog section * Stop re-merging the Codex plan-utilization history with itself on every refresh (steipete#3141) `materializeCodexPlanUtilizationHistoryIfNeeded` exists to fold legacy, opaque and unscoped Codex plan-utilization buckets into the canonical account bucket. Its scoped loop also appended the canonical bucket's own histories to `historiesToMerge` — `matchesTargetContinuity` is true for `rawKey == canonicalKey`, and only the removal of the old key was guarded — so `guard !historiesToMerge.isEmpty` never fired once the canonical bucket had any history, and the migration merge ran on every successful provider refresh and every menu open, merging the history with itself. That merge is quadratic: `updatedPlanUtilizationEntries` copied the whole entry array per entry, scanned it linearly for the insertion point, and allocated the same-hour slice. Measured with an optimized standalone reproduction over a real three-month-old history (session 1909 entries, weekly 2239): 20.6 ms of MainActor time per call, scaling ~3.9x per doubling. `planUtilizationMaxSamples` allows 17520 entries per series, so it would keep growing. Two changes: - Track whether a foreign source actually contributed and require that in the guard, so the canonical-only case returns without merging or rewriting anything. Every path where a legacy, opaque or unscoped bucket contributes is untouched; `legacyRawKeysToRemove` is populated only in branches that also set the flag, so no removal is skipped, and `providerBuckets.unscoped` is cleared only inside the branch that sets it. - Make the merge itself near-linear: `updatedPlanUtilizationEntries` mutates the array in place and finds the insertion point with a binary search for the same strict upper bound (with a fast path for the common append), and `mergedPlanUtilizationHistories` accumulates per series and builds each history once. The binary search assumes entries are sorted by `capturedAt`, which every in-app producer guaranteed through `PlanUtilizationSeriesHistory`'s designated initializer — except the synthesized `Codable` decoder, which assigned entries verbatim from JSON. An explicit `init(from:)` now routes decoding through that initializer, so an on-disk history written by an older build or edited by hand cannot smuggle in an unsorted series. The skipped self-merge also incidentally re-canonicalized per-hour peaks on read; that repair belongs at load time, not on every refresh, and is not reintroduced here. The visible effect is that at most one extra real observation per affected hour is kept. Tests: canonical-only history is returned untouched and enqueues no persistence write (the history revision is unchanged); a genuine foreign merge matches an explicit expected result across overlapping hours, out-of-order sources, distinct series and retention trimming; the binary search's upper-bound contract is pinned directly through a DEBUG shim over an array with a run of equal timestamps (a lower bound would return a different index); and decoding a series whose JSON entries are out of order yields a sorted series. Implemented by grok-4.6 (xhigh) via implementation-loop; reviewed hunk by hunk plus an independent deep review that confirmed both equivalences by differential fuzzing (200k sorted cases with no mismatch) and found the decoder gap, fixed in one iterate round. Gatekeeper line anchors for the touched file were re-verified independently. The DEBUG sortedness assertion is checked once per merged series rather than once per inserted entry: a per-entry check is itself O(n) and reintroduced, in debug builds, exactly the quadratic scan this insertion path removes (measured over the real 4160-entry history: a legacy migration took ~1000 ms with the per-entry assertion versus ~15 ms without it). Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * Fix Codex day cost blanked by trace-only priority turns (steipete#3150) Row ownership evidence compared the retained rows against the persisted standard/priority split using the trace database's tier classification. The persisted maps come from the rows' own pricingMode, so a turn the trace reports as priority after its rows were persisted as standard read as a row-ownership mismatch, the rows lost trust, and the day fell back to the aggregate — which returns nil for long-context tiered models, so the whole day's cost disappeared from the menu, the chart and the window total. Judge retention against both classifications and flag only a group that matches neither. A wrongly retained row set still fails both, because the persisted totals are canonical for the file and tier classification never changes how many tokens the rows carry. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * docs: credit steipete#3141 and steipete#3150 changelog entries * fix(qwen-cloud): restore Brave browser support in cookie import (steipete#3148) * fix(qwen-cloud): restore Brave browser support, narrowed to Chrome+Brave per AGENTS.md Qwen Cloud's cookie import was restricted to [.chrome] only (commit 529cc6c 'Keep Qwen imports Chrome-only'). Brave users hit 'No Qwen Cloud session cookies found in browsers' even when they had a valid Qwen Cloud session in Brave, because their cookies were never probed. This commit restores Brave in the import order, but follows AGENTS.md L48 ('default Chrome-only when possible to avoid other browser prompts; override via browser list when needed'). The override is the minimum necessary: Chrome + Brave. The other Chromium browsers (chromeBeta, edge, arc, firefox, safari) are deliberately omitted to avoid unsolicited Keychain / browser-store access prompts on automatic refreshes from browsers that don't carry a Qwen Cloud session. Brave is kept because it shares the same Chromium Safe Storage format as Chrome and is a common Qwen Cloud authentication target. Also adds docs/qwen-cloud-proof/README.md with the redacted end-to-end proof captured against the live Qwen Cloud API from the user's Mac after granting the modified binary access to 'Brave Safe Storage' in macOS Keychain. * fix(qwen-cloud): recovery message now names Brave alongside Chrome ClawSweeper P2 follow-up on steipete#3148: when the Brave cookie import fails, QwenCloudSettingsError.missingCookie's recovery message still told users to sign in to Chrome and grant access to Chrome Safe Storage. Now that Brave is a supported source, the message must name both browsers and their respective Safe Storage entries, otherwise a Brave-only user would be told to use Chrome and never find the working path. Updates the error description to: 'No Qwen Cloud session cookies found in browsers. Sign in to Qwen Cloud in Chrome or Brave, allow CodexBar to access the corresponding Safe Storage in Keychain Access (Chrome Safe Storage and/or Brave Safe Storage), or paste a manual Cookie header.' Adds focused test coverage: - missing cookie error mentions both supported browsers and their safe storage - missing cookie error appends non-empty details - missing cookie error omits empty details 35/35 Qwen Cloud tests pass (32 prior + 3 new). * Fix OpenRouter completed-day activity query (steipete#3138) * Preserve unknown Grok period usage (steipete#3159) Co-authored-by: anupamchugh <8416306+anupamchugh@users.noreply.github.com> * fix: report non-writable CLI path conflicts (steipete#3153) * fix: prefer successful CLI install status * fix: keep CLI path conflicts visible * fix: report non-writable CLI path conflicts * docs: add CLI conflict behavior proof * docs: add CLI install comparison screenshots * Fix single-quota icon scaling (steipete#3155) * docs: credit steipete#3138 steipete#3148 steipete#3153 steipete#3155 steipete#3159 changelog entries * fix(spend): silent refresh and invalidation coverage (steipete#3106) * fix(spend): bucket calendar for all heatmap dates and full revision hash - SpendActivityDateFormatting.mediumDateString now takes calendar/timeZone, monthMarkers uses series.calendar, tooltips and accessibility use bucket calendar. - selectedDay renormalized on calendar change to keep toggle correct. - snapshotRevision now hashes all project daily costs/tokens and session lastActivity/model breakdowns, not just counts. Fixes ClawSweeper P2 for steipete#3106. * fix(gatekeeper): update anchors and add provider-specific design markers for spend dashboard * Update provider gatekeeper anchors for v0.54 rebase * fix(test): pin claude spend snapshot in observation test * fix(test): seed pinned claude spend publication before first snapshot * fix: resolve remaining conflict markers from gatekeeper rebase * fix(lint): shorten Sakana test lines * fix(spend): restore heatmap calendar property lost in rebase * Extend spend publication test wait * Restore spend gatekeeper anchors after rebase * fix(spend): sync independent snapshot and bucket calendar normalization for 3106 - publishSpendDashboardTokenSnapshotState now calls synchronizeSharedSpendDashboardAfterTokenPublication - heatmap calendar onChange no longer renormalizes selectedDay via stale controller - SpendDashboardController.update now normalizes selectedDay atomically when bucketTimeZoneIdentifier changes - update gatekeeper anchors for shifted lines (1620,1649,1666,1693) * test(spend): cover independent snapshot sync for 3106 Exercise the direct independent publication path added at UsageStore+SpendDashboardTokenCost.swift:181. The prior focused test seeded Claude before observation and then used the regular Codex publisher, which already syncs independently, so removing that line would not fail. Add a post-start Claude snapshot via _setSpendDashboardTokenSnapshotForTesting and assert the shared dashboard debounced sync is scheduled and the publication inputs update. Verified: swiftformat clean, swiftlint --strict clean, swift test --filter SpendDashboardPublicationTests 18 tests passed. * docs: add fresh-bundle proof for 3106 Add redacted menu-icon crop and dashboard snapshot from debug build 2798eec (swift build --target CodexBarCLI, .build/debug/CodexBarCLI dashboard --pretty). The snapshot shows the shared spend controller produces a dashboard with provider rows/windows, confirming the independent-sync and calendar paths are live in the fresh binary. * docs: add menu and Spend dashboard screenshots for 3106 Add redacted screenshots from fresh debug build 45ba984: - 3106-menu-after-fix.png: menu bar extra open, showing provider rows - 3106-settings-after-fix.png: Settings window (general) - 3106-spend-dashboard-after-fix.png: Usage & Spend pane (usageSpend) with heatmap and Overview, confirming the shared controller renders in the fresh bundle. * docs: remove screenshots for 3106 per request Keep only the redacted CLI dashboard snapshot JSON as fresh-bundle proof; screenshots are not needed. * Improve Antigravity retrieval: retired Flash alias and offline fallback (steipete#3119) * fix(antigravity): allow OAuth errors to fallback to offline when local data exists Fix P2 from Codex review on steipete#3119: AntigravityOAuthFetchStrategy.shouldFallback now checks hasOfflineData, so expired credentials do not block offline. * fix(antigravity): unbind offline account, read app-data, bound scans + proof - Offline snapshot now has nil accountEmail (P1) - OfflineStore also counts $HOME/.gemini/antigravity and .../conversations (P2) - SpendDashboardController bounds Codex scans to 3 concurrent (P2) - Add AntigravityOfflineFallbackProofTests covering app-data and nil email * fix: remove broken proof test, keep P1/P2 fixes and shell proof * fix(gatekeeper): update SpendDashboardController anchors after bounding Codex scans * fix: revert bounded Codex scans (keep offline P1/P2), restore gatekeeper * feat(spend): add tokscale-compatible local readers for Cursor and Antigravity (steipete#3113) * feat(spend): add tokscale-compatible local readers for Cursor and Antigravity - Cursor: read ~/.config/tokscale/cursor-cache/usage*.csv (v1/v2/v3) with tokstyle column handling, cacheWrite = with-without, noon UTC for date-only, and CostUsageDailyReport aggregation. (Sources/CodexBarCore/Providers/Cursor/CursorLocalCSVReader.swift:1) - Antigravity: read ~/.config/tokscale/antigravity-cache/sessions/*.jsonl (tokscale JSONL) and stub for ~/.gemini/antigravity-cli/*.db direct SQLite (ProtoReader to follow). Handles session_meta fallback and dedup. (Sources/CodexBarCore/Providers/Antigravity/AntigravityLocalReader.swift:1) - CostUsageFetcher: local fallback before remote for Cursor (offline) and primary for Antigravity (quota-only before), with Provider-specific by design comments for gatekeeper. (Sources/CodexBarCore/CostUsageFetcher.swift:440) - Antigravity descriptor: enable supportsTokenSnapshot for spend dashboard. (Sources/CodexBarCore/Providers/Antigravity/AntigravityProviderDescriptor.swift:51) Reproduced from /tmp/opencodex/src/adapters/cursor/protobuf-events.ts:218 and /tmp/tokscale/crates/tokscale-core/src/sessions/{cursor,antigravity_cli}.rs Phase 1 of opencodex/tokscale plan, offline-first, no auth. * test(readers): cover cursor csv schemas and antigravity cache fallback * fix(test): include antigravity in cost capable dashboard sources * fix(spend): honor CSV total tokens and add Antigravity Linux capability * fix(test): honor cursor CSV total tokens column in aggregation * fix(spend): repair 3113 tokscale readers P1s - catch remote Cursor errors before falling back to local CSV - recompute summaries after window filtering for Cursor and Antigravity - keep Antigravity costs nil (unpriced) and deduplicate by responseId - parse date-only CSV rows with UTC calendar - thread fallback calendar through loaders * fix(lint): repair 3113 build and format - calendar before now in makeDailyReport - implicit optional init - wrap long lines and andOperator * style: swiftformat wrap for 3113 * Fix 3113 provider gatekeeper anchors * fix(spend): address 3113 review findings -- freshness, calendar, date-only, fixture model - Preserve cache freshness: return nil when filtered window is empty instead of publishing established zero with now timestamp - Pass pinned calendar into tokenSnapshot for Cursor/Antigravity local snapshots - Keep date-only Cursor CSV rows in configured calendar's noon, not UTC noon - Use clearly fictitious test model test-model-antigravity-a * style: fix line length for fixture model * fix(test): update gatekeeper anchors for CostUsageFetcher line drift Allowlist lines 1339->1335 and 1695->1691 after 075eac7 freshness/calendar fixes * test: repair gatekeeper anchors and regenerate parser hash on merged tree --------- Co-authored-by: Yuxin-Qiao <2242016570@qq.com> Co-authored-by: Peter Steinberger <steipete@gmail.com> Co-authored-by: Olddonkey <olddonkeyblog@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Umut Keltek <35880258+umutkeltek@users.noreply.github.com> Co-authored-by: kiranmagic7 <kiranmagic@proton.me> Co-authored-by: Anupam Chugh <anupam.chugh@gmail.com> Co-authored-by: anupamchugh <8416306+anupamchugh@users.noreply.github.com> Co-authored-by: yicone <yicone@gmail.com> Co-authored-by: Akshay Prabhu <12824090+akshayprabhu200@users.noreply.github.com>
|
#3141 and #3150 are merged — thanks. This PR (and #3140/#3135 stacked above it) now conflicts with the landed parser-hash/scanner changes from #3150 and #3120. Please rebase onto current main: regenerate the parser hash on the merged tree ( |
# Conflicts: # CHANGELOG.md # Sources/CodexBarCore/Generated/CodexParserHash.generated.swift # Sources/CodexBarCore/Vendored/CostUsage/CostUsageStore.swift
Summary
The OpenCodex spend source (
~/.opencodex/usage.jsonl→OpenCodexUsageFanOut→OpenCodexUsageAggregator.snapshot) re-resolved its pricing context per entry on every refresh:listPriceUSDcalledCostUsagePricing.codexCostUSDwithout a pre-resolved models.dev catalog, so each call went throughModelsDevCache.load→FileManager.attributesOfItem— astatplus an extended-attribute read — and without a pre-resolved custom-pricing overlay, so each call also re-resolved the overlay file location.Calendarper entry.On a 35k-entry log (all inside the 30-day window) that is ~100k stat+xattr syscalls and ~70k Calendar interval computations per refresh. In the running app this was the 25–35 s CPU spike on every adaptive refresh (sampled on 0.54.0:
snapshotsBySubscription→attributesOfItem→getxattr/listxattr, plus Calendar/ICU and String hashing).Changes (snapshot output is byte-identical)
OpenCodexUsageAggregator,OpenCodexUsageFanOut)[start, end)intervals (DST-correct; no 86400/3600 arithmetic). Day keys still come fromCostUsageLocalDay.key, which derives y-m-d from the same Gregorian-in-timezone calendar whose day interval the memo caches, so the memo is exact.ModelsDevCache.loadreads (mtime, size) via POSIXstatinstead ofattributesOfItem(which also reads xattrs); memo/invalidation semantics unchanged. This helps every caller repo-wide. (ModelsDevPricing.swift)CodexParserHashis regenerated becauseModelsDevPricing.swiftis in the hashed set; the previous hash (3c984b655688593f) is added tocompatiblePredecessorParserHashessince parsing and the persisted row shape are unchanged, so existingcost-usage.sqlitestores are adopted on upgrade instead of rebuilt (same pattern as Persist Codex priority-turn scan cursor across relaunches #3130).Not in this PR (follow-up):
OpenCodexUsageStorekeys its sqlite cache onpath|size|mtime, so the append-only log is fully re-parsed (and DELETE+re-inserted) whenever it grew. That is where the remaining transient memory on this path lives; making the store incremental is a separate change.Measured
Release
CodexBarCLI, isolated cache root, same machine, real 41.7 MB / ~35k-entryusage.jsonl(all inside the window),cost --provider codex --format json --days 30.27c7f334eCLI JSON output is identical in both scenarios (diffed with only
updatedAtstripped). Fullcost --provider codexrun on live data in the app's steady-state case (log grew since the last refresh): ~11 s → ~3.5 s real (7.3–9.2 s → 3.2 s user); cold 26 s → 14 s. Peak footprint is unchanged — that is the re-parse path noted above.Tests
OpenCodexUsageFanOutTests: equivalence of the newsnapshotagainst an independent reference implementation (mixed providers, reported/estimated/unreported/unsupported entries, custom-pricing overlay, duplicate request IDs, entries outside the window, DST transitions in America/Los_Angeles, pricing through a fixture models.dev catalog so the catalog path is actually exercised); metadata-read counting proving one catalog load per snapshot for 60 entries and zero with an injected catalog; day/hour memo boundary cases (spring-forward gap, fall-back repeated hour, exact hour/midnight boundaries, and America/Santiago midnight DST transitions) compared againstCostUsageLocalDay.key/calendar.dateInterval(of: .hour)computed per entry.ModelsDevPricingTests: memo still invalidates on cache-file size change and on mtime change after thestatswitch; exactly one metadata read perModelsDevCache.load(task-local recorder).CostUsageStoreTests: predecessor-hash set includes3c984b655688593f.Process
Implemented by grok-4.6 (xhigh) via implementation-loop; reviewed hunk by hunk plus an independent deep review (no correctness defect; its findings — overlay-once, empty-catalog substitution, task-local recorder, deterministic catalog-path test, Santiago DST cases — were applied in one iterate round). Full
make checkandmake testrun locally on the final tree.🤖 Generated with Claude Code