feat(api): configurable unpriced-usage budget policy + admin notice (#14604) - #14608
xiaoyaner0201 wants to merge 6 commits into
Conversation
64c4704 to
43822f0
Compare
…iegosouzapw#14604) Per-key USD limits fail closed on usage with no pricing row since diegosouzapw#13257 (diegosouzapw#12341). Keep that as the default and make the choice explicit: - UNPRICED_USAGE_BUDGET_POLICY flag: fail_closed (default, unchanged verdict) or count_as_zero (unpriced usage counts as $0, priced spend is still enforced). Unknown values and flag-store failures fall back to fail_closed. A cost lookup that throws is not "unpriced": it still blocks under every policy (calculateCostDetailed now reports failed: true) and the rejection says the cost could not be calculated instead of blaming missing prices. - When a key is blocked only by unpriced usage, the 400 names the unpriced models instead of "reached its weekly usage quota (3%)". If that window has no fixed reset time (rolling weekly window), the reset clause is left out instead of printing "Resets in unknown". A priced overage in any window keeps the regular quota message, reporting the window that is actually over on priced spend. - GET /api/pricing/unpriced-usage (management auth) and a dashboard home banner list unpriced provider/model pairs from the last 8 days, the keys with an enforceable limit they affect, and link to the pricing editor. - i18n: the banner, the flag's label/description and its two enum values are next-intl keys in en.json, translated into every UI locale (ICU plurals, protected tokens kept verbatim). The banner no longer carries hard-coded English fallbacks.
43822f0 to
6df78f4
Compare
CI triage for 6df78f4I compared every red gate against the exact base Fixed in this push (caused by this PR)
Also red on the base (same output on both trees)
I did not touch any of those files or baselines to force green. One more note on the first run for |
|
Thanks @xiaoyaner0201, this is a careful follow-up to #12341. I trial-merged it together with #14603 and #14619: the only conflict is the import block with #14603 (keep both), and the combined focused suites pass (108/108). All 67 catalogs carry the new keys with no |
|
CI follow-up on head
The PR is not CI-green. I have not changed unrelated baseline files or weakened the gate to make it green; this comment records the exact scope of the differential evidence for maintainer triage. |
Maintainer sequencing follow-upThe candidate at The remaining PR red checks are not introduced by this branch:
I have not changed unrelated baseline files or weakened gates to manufacture green CI. Please sequence #14608 after the related usage PRs as appropriate (it currently carries |
|
Re-homed to |
Summary
Closes #14604. Follow-up to #12341 / #13257.
fail_closedremains the default. This adds an explicit operator choice, actionable rejection messages and a management-authenticated notice for missing model prices.UNPRICED_USAGE_BUDGET_POLICYacceptsfail_closedorcount_as_zero. Unknown values or an unreadable flag store fall back tofail_closed.count_as_zeropermits genuinely unpriced usage while enforcing limits on priced spend. A pricing lookup that throws remains fail-closed under both policies;calculateCostDetailed(...).faileddistinguishes that case from a missing price.Retry-After, rather than presenting a repairable pricing problem as a priced quota overage. Genuine priced overages retain the release's non-Anthropic 429 contract; Anthropic messages retain their existing 400 behavior. When daily and weekly windows overlap, the quota response and reset metadata follow the window actually over on priced spend.GET /api/pricing/unpriced-usagerequires management auth. It aggregates the last eight days of unpriced provider/model usage, request counts and affected keys, and reports the active policy.fail_closed, affected limited keys show a red alert; if no limited key is affected, the banner stays amber/status rather than claiming a zero-key blockage.count_as_zerouses amber/status. Dismissal lasts only for the current page view.dailyUnpricedModels,weeklyUnpricedModels,dailyPricingFailure,weeklyPricingFailureandunpricedUsagePolicy.Review and landing status
The existing
deferred-v3.8.52label is preserved. This remains a dedicated, one-to-one review slice, not an automatic merge-train candidate and not a claim of maintainer approval or full green CI. The public base remainsrelease/v3.8.51.The 67 locale catalogs are most of the 83-file footprint. Runtime policy, its report, UI and conformance tests form one buildable slice; separating them into independently published partial states would leave the operator contract incomplete. No next backend or unrelated feature is included.
Exact latest-base refresh — 2026-09-26
c1854ca94a1b6a139278dbde92ca21c21c401798.199fc173bbb67eead9c2e5eee1e15d33e2cdb999.6c766317bcd07e92784459473a33404a64ce8372.ae2ba35852d4e5a55486a1c0e6a779105564fd6d.Conflict reconciliation preserved release changes rather than restoring old inline gateway behavior. Independent review first rejected the candidate's generic 429 treatment of missing-price blocks, zero-affected-key banner severity and description-key wiring. Those findings were repaired with tests-first RED→GREEN. The final read-only Claude Code / Opus review verified the actual commit, tree and both parents, production consumers, management auth, error-body/header behavior, all 67 catalog keys and the previously uncertain findings: PASS, zero confirmed P0/P1/P2. This is an internal pre-push review, not a GitHub maintainer review.
Verification for the exact tree
Focused tests
91 passed, 0 failed. The same three pre-existing suites, with the identical Node command, pass 80/80 on the exact release base; the fourth file is this PR's new policy coverage. Tests exercise real usage rows, pricing lookup and enforcement, including lookup exceptions and mixed daily/weekly reset selection.
2 passed, 0 failed. Both the initial review-rejected core cases and the zero-affected-key banner case were demonstrated failing before their corresponding repair.
Fresh final runs passed:
npm run typecheck:corenpm run check:complexity-ratchets -- --base-ref=ae2ba35852d4e5a55486a1c0e6a779105564fd6dnpm run check:file-sizenpm run check:cyclesnpm run check:pr-test-policy -- --base-ref=ae2ba35852d4e5a55486a1c0e6a779105564fd6dnpm run i18n:check-ui-coveragegit diff --checkagainst the exact release base.The repository pre-commit checks and targeted ESLint passed before the commit was frozen. No code changed after the final source-aware independent review.
Full-suite comparison — explicitly not all green
All four CI shard commands were run on the candidate and the exact
ae2ba358base, with the same command for each shard:Both sides have failures. Comparing the aggregate failing test-name sets yields no candidate-only failing test names. These include environment-sensitive subprocess/network tests and existing i18n/documentation contracts; they were not edited to force green. The full UI Vitest run likewise has the same 21 failing test names on both trees, while this PR's new two-case banner suite passes.
The four-shard run preceded final Prettier-only changes in four files. Their whitespace-minified executable JS outputs were checked byte-for-byte identical before/after; the final exact tree then reran the 91 core and 2 UI tests successfully. The full four shards were not rerun after formatting, and are not presented as an exact post-format full-green receipt.
check:env-doc-sync,check:docs-all, fullnpm run lint,npm run i18n:checkand changelog integrity remain red with matching signatures on the exact base and candidate. The repository's canonical-temp-directory setting was used for the complexity comparison; neither its implementation nor its baseline was changed.Fresh remote checks on
c1854ca94ahave completed in run 36226363250. API Route Typecheck, ESLint, Vitest, Semgrep, classification and merge integrity pass. Docs Gates, Fast Quality Gates and all four unit shards fail; the advisory build is skipped, and the Mergify checks are neutral. All 32 failed remote unit-test names have corresponding failures in the exact-base four-shard results. The two Fast Quality failures were also rerun with the identical commands on both trees: strict mutation coverage reports the same ten missing covering tests (only the scanned test count differs), and dashboard typecheck reports the identicalNoAuthAccountCard.tsxTS2345. These are inherited base failures, not a full-green CI claim.Screenshots and previous live evidence
These screenshots show the original implementation, before the latest-base review repair; they are not fresh browser verification of this head:
The previous live HTTP/browser check used a throwaway data directory and seeded usage rows, and is historical evidence only. This refresh did not start a server, touch a running instance or use production data.
Risks / boundaries