Skip to content

feat(api): configurable unpriced-usage budget policy + admin notice (#14604) - #14608

Open
xiaoyaner0201 wants to merge 6 commits into
diegosouzapw:release/v3.8.52from
xiaoyaner0201:fix/unpriced-usage-policy
Open

xiaoyaner0201 wants to merge 6 commits into
diegosouzapw:release/v3.8.52from
xiaoyaner0201:fix/unpriced-usage-policy

Conversation

@xiaoyaner0201

@xiaoyaner0201 xiaoyaner0201 commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Closes #14604. Follow-up to #12341 / #13257.

fail_closed remains the default. This adds an explicit operator choice, actionable rejection messages and a management-authenticated notice for missing model prices.

  • UNPRICED_USAGE_BUDGET_POLICY accepts fail_closed or count_as_zero. Unknown values or an unreadable flag store fall back to fail_closed.
  • count_as_zero permits genuinely unpriced usage while enforcing limits on priced spend. A pricing lookup that throws remains fail-closed under both policies; calculateCostDetailed(...).failed distinguishes that case from a missing price.
  • A missing-price-only or lookup-failure rejection is 400 without Retry-After, rather than presenting a repairable pricing problem as a priced quota overage. Genuine priced overages retain the release's non-Anthropic 429 contract; Anthropic messages retain their existing 400 behavior. When daily and weekly windows overlap, the quota response and reset metadata follow the window actually over on priced spend.
  • GET /api/pricing/unpriced-usage requires management auth. It aggregates the last eight days of unpriced provider/model usage, request counts and affected keys, and reports the active policy.
  • The dashboard home banner links to pricing configuration. Under fail_closed, affected limited keys show a red alert; if no limited key is affected, the banner stays amber/status rather than claiming a zero-key blockage. count_as_zero uses amber/status. Dismissal lasts only for the current page view.
  • Additive status fields: dailyUnpricedModels, weeklyUnpricedModels, dailyPricingFailure, weeklyPricingFailure and unpricedUsagePolicy.
  • The banner's seven keys and the policy definition/enum keys exist in all 67 UI catalogs. The flag definition uses the real nested description key; ICU argument names are consistent across catalogs.

Review and landing status

The existing deferred-v3.8.52 label is preserved. This remains a dedicated, one-to-one review slice, not an automatic merge-train candidate and not a claim of maintainer approval or full green CI. The public base remains release/v3.8.51.

The 67 locale catalogs are most of the 83-file footprint. Runtime policy, its report, UI and conformance tests form one buildable slice; separating them into independently published partial states would leave the operator contract incomplete. No next backend or unrelated feature is included.

Exact latest-base refresh — 2026-09-26

  • Published head: c1854ca94a1b6a139278dbde92ca21c21c401798.
  • Exact tree: 199fc173bbb67eead9c2e5eee1e15d33e2cdb999.
  • First parent: prior published PR head 6c766317bcd07e92784459473a33404a64ce8372.
  • Second parent / tested release base: ae2ba35852d4e5a55486a1c0e6a779105564fd6d.
  • Ordinary fast-forward push; no rebase, amend, force-push or Draft/Ready transition.
  • Against that release base: 83 files, +1936/-181, including 67 locale catalogs. The original 82-file scope is retained, with one focused banner test file added during review repair.
  • Both release feature flags and this policy flag are preserved. The flag catalog/header count is 81.

Conflict reconciliation preserved release changes rather than restoring old inline gateway behavior. Independent review first rejected the candidate's generic 429 treatment of missing-price blocks, zero-affected-key banner severity and description-key wiring. Those findings were repaired with tests-first RED→GREEN. The final read-only Claude Code / Opus review verified the actual commit, tree and both parents, production consumers, management auth, error-body/header behavior, all 67 catalog keys and the previously uncertain findings: PASS, zero confirmed P0/P1/P2. This is an internal pre-push review, not a GitHub maintainer review.

Verification for the exact tree

Focused tests

node --import tsx/esm --test \
  tests/unit/api-key-usage-limits.test.ts \
  tests/unit/feature-flags-settings.test.ts \
  tests/unit/server-owned-tool-loop-flag.test.ts \
  tests/unit/unpriced-usage-budget-policy-12341.test.ts

91 passed, 0 failed. The same three pre-existing suites, with the identical Node command, pass 80/80 on the exact release base; the fourth file is this PR's new policy coverage. Tests exercise real usage rows, pricing lookup and enforcement, including lookup exceptions and mixed daily/weekly reset selection.

./node_modules/.bin/vitest run tests/unit/ui/unpriced-usage-banner.test.tsx

2 passed, 0 failed. Both the initial review-rejected core cases and the zero-affected-key banner case were demonstrated failing before their corresponding repair.

Fresh final runs passed:

  • npm run typecheck:core
  • npm run check:complexity-ratchets -- --base-ref=ae2ba35852d4e5a55486a1c0e6a779105564fd6d
  • npm run check:file-size
  • npm run check:cycles
  • npm run check:pr-test-policy -- --base-ref=ae2ba35852d4e5a55486a1c0e6a779105564fd6d
  • npm run i18n:check-ui-coverage
  • git diff --check against the exact release base.

The repository pre-commit checks and targeted ESLint passed before the commit was frozen. No code changed after the final source-aware independent review.

Full-suite comparison — explicitly not all green

All four CI shard commands were run on the candidate and the exact ae2ba358 base, with the same command for each shard:

TEST_SHARD=1/4 npm run test:unit:ci:shard
TEST_SHARD=2/4 npm run test:unit:ci:shard
TEST_SHARD=3/4 npm run test:unit:ci:shard
TEST_SHARD=4/4 npm run test:unit:ci:shard

Both sides have failures. Comparing the aggregate failing test-name sets yields no candidate-only failing test names. These include environment-sensitive subprocess/network tests and existing i18n/documentation contracts; they were not edited to force green. The full UI Vitest run likewise has the same 21 failing test names on both trees, while this PR's new two-case banner suite passes.

The four-shard run preceded final Prettier-only changes in four files. Their whitespace-minified executable JS outputs were checked byte-for-byte identical before/after; the final exact tree then reran the 91 core and 2 UI tests successfully. The full four shards were not rerun after formatting, and are not presented as an exact post-format full-green receipt.

check:env-doc-sync, check:docs-all, full npm run lint, npm run i18n:check and changelog integrity remain red with matching signatures on the exact base and candidate. The repository's canonical-temp-directory setting was used for the complexity comparison; neither its implementation nor its baseline was changed.

Fresh remote checks on c1854ca94a have completed in run 36226363250. API Route Typecheck, ESLint, Vitest, Semgrep, classification and merge integrity pass. Docs Gates, Fast Quality Gates and all four unit shards fail; the advisory build is skipped, and the Mergify checks are neutral. All 32 failed remote unit-test names have corresponding failures in the exact-base four-shard results. The two Fast Quality failures were also rerun with the identical commands on both trees: strict mutation coverage reports the same ten missing covering tests (only the scanned test count differs), and dashboard typecheck reports the identical NoAuthAccountCard.tsx TS2345. These are inherited base failures, not a full-green CI claim.

Screenshots and previous live evidence

These screenshots show the original implementation, before the latest-base review repair; they are not fresh browser verification of this head:

The previous live HTTP/browser check used a throwaway data directory and seeded usage rows, and is historical evidence only. This refresh did not start a server, touch a running instance or use production data.

Risks / boundaries

@xiaoyaner0201
xiaoyaner0201 force-pushed the fix/unpriced-usage-policy branch from 64c4704 to 43822f0 Compare September 23, 2026 08:25
@xiaoyaner0201
xiaoyaner0201 marked this pull request as ready for review September 23, 2026 08:26
@xiaoyaner0201
xiaoyaner0201 marked this pull request as draft September 23, 2026 08:34
@xiaoyaner0201
xiaoyaner0201 marked this pull request as ready for review September 23, 2026 08:34
…iegosouzapw#14604)

Per-key USD limits fail closed on usage with no pricing row since diegosouzapw#13257
(diegosouzapw#12341). Keep that as the default and make the choice explicit:

- UNPRICED_USAGE_BUDGET_POLICY flag: fail_closed (default, unchanged
  verdict) or count_as_zero (unpriced usage counts as $0, priced spend is
  still enforced). Unknown values and flag-store failures fall back to
  fail_closed. A cost lookup that throws is not "unpriced": it still blocks
  under every policy (calculateCostDetailed now reports failed: true) and
  the rejection says the cost could not be calculated instead of blaming
  missing prices.
- When a key is blocked only by unpriced usage, the 400 names the unpriced
  models instead of "reached its weekly usage quota (3%)". If that window
  has no fixed reset time (rolling weekly window), the reset clause is
  left out instead of printing "Resets in unknown". A priced overage
  in any window keeps the regular quota message, reporting the window that
  is actually over on priced spend.
- GET /api/pricing/unpriced-usage (management auth) and a dashboard home
  banner list unpriced provider/model pairs from the last 8 days, the keys
  with an enforceable limit they affect, and link to the pricing editor.
- i18n: the banner, the flag's label/description and its two enum values
  are next-intl keys in en.json, translated into every UI locale (ICU
  plurals, protected tokens kept verbatim). The banner no longer carries
  hard-coded English fallbacks.
@xiaoyaner0201
xiaoyaner0201 force-pushed the fix/unpriced-usage-policy branch from 43822f0 to 6df78f4 Compare September 23, 2026 09:01
@xiaoyaner0201

Copy link
Copy Markdown
Contributor Author

CI triage for 6df78f4

I compared every red gate against the exact base 18bbb1019. Each gate below was run with the identical command on a clean worktree of the base and on this branch, and the outputs were diffed.

Fixed in this push (caused by this PR)

  • check:complexity-ratchets: buildUsageLimitExceededMessage reached cyclomatic 17 when I added the "report the window that is actually over on priced spend" selection. I moved that selection into two private helpers, isPricedOverLimit and shouldReportDailyWindow. The file is back to its two pre-existing violations, findWeeklyQuotaResetAt and getProviderWeeklyWindow, both untouched. Behavior is unchanged: I checked the full truth table against the old inline expression, the focused suites pass 19/19, and a sabotage run of the helper fails a priced weekly overage is reported even when the daily window is blocked only by unpriced usage.

Also red on the base (same output on both trees)

  • API Route Typecheck and check:open-sse-typecheck: open-sse/executors/auggie.ts TS2769/TS18047, and src/app/api/v1/combos/projectCombo.ts TS2459/TS2724. Both trees report apiTypecheckErrors=294.
  • check:dashboard-typecheck: src/lib/combos/intelligentRouting.ts TS2698.
  • typecheck:core: src/lib/services/cliproxyAccountHealth.ts:157 TS2322.
  • check:deps: @opencode/plugin is not in the allowlist.
  • check:env-doc-sync (Docs Gates): DEEP_HEALTH_CHECK_ENABLED is missing from .env.example.
  • check:mutation-test-coverage: the same 12 test files are missing from stryker.conf.json, all for auth, circuit-breaker and combo modules. The only output difference is the scanned count, 5537 vs 5538, which is the new test file this PR adds.
  • Unit Tests fast-path: 32 failures across 24 files, covering chatCore, compression, combos and provider catalogs. I re-ran those 24 files on both trees with the same command. Both trees give 385 tests, 353 pass and 32 fail, with identical sorted failing names. Every CI failure reproduced locally, and none of them fails only on this branch.

I did not touch any of those files or baselines to force green.

One more note on the first run for 43822f098: the draft-to-ready switch fired two runs close together, and cancel-in-progress cancelled the full one. The surviving run skipped the heavy jobs, so the first rollup showed a "success" that tested nothing. I toggled the PR back to draft and to ready again to get a real run.

@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks @xiaoyaner0201, this is a careful follow-up to #12341. fail_closed stays the default, lookup failures are kept separate from missing prices, and the new message finally tells operators which model is unpriced.

I trial-merged it together with #14603 and #14619: the only conflict is the import block with #14603 (keep both), and the combined focused suites pass (108/108). All 67 catalogs carry the new keys with no __MISSING__ markers. The red CI checks are inherited base reds (auggie typecheck #14496, DEEP_HEALTH_CHECK_ENABLED, @opencode/plugin), identical on #14603. We plan to land this one first.

@diegosouzapw diegosouzapw added the deferred-v3.8.52 Grande demais / suspeito para o lote atual; precisa de sessão dedicada no ciclo v3.8.52 label Sep 25, 2026
@diegosouzapw diegosouzapw changed the title feat(api): configurable unpriced-usage budget policy + admin notice (#14604) [defer] feat(api): configurable unpriced-usage budget policy + admin notice (#14604) Sep 25, 2026
@xiaoyaner0201

Copy link
Copy Markdown
Contributor Author

CI follow-up on head 30f940fea9 against exact release base a58000c768:

  • The previously reported dual-window reset issue is fixed: when both daily and weekly limits block a key, the response no longer promises the earlier reset as recovery or sends retry_after, reset_at, or Retry-After. On the new head, the focused policy/API-key suites pass (93/93), UI tests pass (2/2), and typecheck:core and check:open-sse-typecheck pass.
  • The current Fast Quality Gates failure is check:dashboard-typecheck: src/shared/components/NoAuthAccountCard.tsx TS2345 (one unbaselined error). I ran the identical gate on a disposable checkout of a58000c768 and on this PR head with the same Node 24.13.0 runtime and dependencies: both exit 1 with dashboardTypecheckErrors=198, the same NoAuthAccountCard.tsx TS2345, and the same 86 baseline entries now absent. The PR does not modify that component or the gate.
  • I also replayed the failing CI test files on base and candidate using the same Node 24.13.0, isolated test data, command, and file lists. For the selected unit files: 46/57 pass on each tree with identical 11 failing test names; 19/21 pass on each tree for MCP bundle/pack-artifact tests with identical two failures; and 69/77 pass on each tree for glossary/Opencode/proxy/status/sidebar tests with identical eight failures. These are measured base failures, not new failures from feat(api): configurable unpriced-usage budget policy + admin notice (#14604) #14608. This is a bounded comparison of the reported failing files, not a claim that the full unit suite passes.
  • Unit shard 2/4 was cancelled during npm run test:unit:ci:shard after 30 minutes and has no final test summary. Its result remains unverified, rather than being counted as a base-matched pass.

The PR is not CI-green. I have not changed unrelated baseline files or weakened the gate to make it green; this comment records the exact scope of the differential evidence for maintainer triage.

@xiaoyaner0201

Copy link
Copy Markdown
Contributor Author

Maintainer sequencing follow-up

The candidate at 30f940fea9 is the reviewed follow-up head and includes the dual-window reset fix. The focused policy/API-key tests are green, UI/typecheck/mutation checks are green locally, and the exact-tree review found no blocking issue.

The remaining PR red checks are not introduced by this branch:

  • check:dashboard-typecheck reports the same src/shared/components/NoAuthAccountCard.tsx TS2345 on a58000c768 and on this head.
  • The reported failing unit files reproduce the same failing test names on the exact base and candidate; shard 2/4 was cancelled and remains unverified.

I have not changed unrelated baseline files or weakened gates to manufacture green CI. Please sequence #14608 after the related usage PRs as appropriate (it currently carries deferred-v3.8.52), or merge it once the inherited base-red checks are addressed. The branch is ready for that maintainer decision.

@diegosouzapw
diegosouzapw changed the base branch from release/v3.8.51 to release/v3.8.52 September 29, 2026 11:21
@diegosouzapw

Copy link
Copy Markdown
Owner

Re-homed to release/v3.8.52: v3.8.51 entered its release freeze, so the branch now belongs to the release captain and development continues on the next cycle. Nothing is wrong with this PR — it just needed a live base. No action needed from you; CI will re-run against the new base.

@diegosouzapw diegosouzapw changed the title [defer] feat(api): configurable unpriced-usage budget policy + admin notice (#14604) feat(api): configurable unpriced-usage budget policy + admin notice (#14604) Oct 1, 2026
@diegosouzapw diegosouzapw removed the deferred-v3.8.52 Grande demais / suspeito para o lote atual; precisa de sessão dedicada no ciclo v3.8.52 label Oct 1, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(backend): one unpriced model blocks a usage-limited API key on every model, with a misleading quota message

2 participants