Repository navigation
feat(sse): add LLM Gateway DevPass quota tracking - #12462
diegosouzapw merged 4 commits into
Conversation
937f7ae to
31ce125
Compare
Surface the LLM Gateway DevPass allowance (GET /v1/key) in OmniRoute's quota telemetry, mirroring the OpenRouter API-key fetcher pattern. - llmgatewayQuotaFetcher.ts: fetch + parse the DevPass /v1/key response (decimal-string USD values), exposing two windows — monthly plan credits and the 7-day premium-model window — with a 45s TTL cache. Pay-as-you-go keys (devPlan "none") and 401/403 fail open (no quota). - Register in chat.ts before registerGenericQuotaFetchers + register the named windows for the dashboard cutoff modal. - usage/llmgateway.ts leaf + usage.ts dispatch case so the Limits page renders the monthly + weekly premium rows. - Add "llmgateway" to USAGE_FETCHER_PROVIDERS, USAGE_SUPPORTED_PROVIDERS, PROVIDER_LIMITS_APIKEY_PROVIDERS, and the dashboard label/order map. - tests: 21 cases covering the parser, auth fail-open, cache TTL, window exhaustion, preflight proceed/block, registration, and the usage leaf.
Move the LLM Gateway fetcher registration out of chat.ts (a frozen file-size-baseline chokepoint) into quotaTrackersBatch.ts, the dedicated side-effect module that exists precisely so new fetchers don't grow chat.ts. The batch import runs at module load, before registerGenericQuotaFetchers(), so the bespoke fetcher still wins over the generic path. Fixes the file-size gate (chat.ts must not grow).
|
Clean quota-tracker addition — mirrors the OpenRouter/Lyceum pattern exactly, 21/21 tests |
31ce125 to
a031e7c
Compare
Resolves the conflict in docs/architecture/CODEBASE_DOCUMENTATION.md — the services table now carries BOTH additions: the base's `requestRejectedStreak.ts` (Resilience) and this branch's `llmgatewayQuotaFetcher.ts` (Quotas). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
|
Thanks @PixmaNts — merging via the release merge-train. Validated in local merge-train /tmp/mt-train3b.log on 192.168.0.113 @ train tip 408e2128791696a18966681953136fc96aeb99b0 (32 PRs boarded): static gates green; full test:unit 40694 tests, 17 failing — every one reproduces on the pure release tip (base-red sweep list), zero new reds. Merged --admin per merge-gates §4/§7. |
01b2467
into
diegosouzapw:release/v3.8.51
Rebased onto release/v3.8.51 after diegosouzapw#12462 (LLM Gateway DevPass quota) merged and the base advanced. The live registry now totals 360 providers, so the auto-generated reference and the diagram/marketing counts are regenerated against the new base: - npm run gen:provider-reference → docs/reference/PROVIDER_REFERENCE.md - 359→360 in README.md, AGENTS.md, llm.txt (provider count), package.json description and the diagram SVGs (readme-hero, promise-pillars, comparison-table, cli-terminal, tier-flow-dark/light) - node scripts/i18n/sync-llm-mirrors.mjs → 65 locale mirrors Addresses the maintainer review asking to rerun gen-provider-reference.ts and push the doc files missing from the original diff, which left Docs Gates red.
Rebased onto release/v3.8.51 after diegosouzapw#12462 (LLM Gateway DevPass quota) merged and the base advanced. The live registry now totals 360 providers, so the auto-generated reference and the diagram/marketing counts are regenerated against the new base: - npm run gen:provider-reference → docs/reference/PROVIDER_REFERENCE.md - 359→360 in README.md, AGENTS.md, llm.txt (provider count), package.json description and the diagram SVGs (readme-hero, promise-pillars, comparison-table, cli-terminal, tier-flow-dark/light) - node scripts/i18n/sync-llm-mirrors.mjs → 65 locale mirrors Addresses the maintainer review asking to rerun gen-provider-reference.ts and push the doc files missing from the original diff, which left Docs Gates red.
…ing (#12474) * feat(providers): add Lyceum pay-per-use provider + credit quota Lyceum (lyceum.technology) is an OpenAI-compatible, usage-based inference provider. Registered as a first-class apikey provider mirroring the llmgateway/openrouter pattern: - registry/lyceum: buildOpenAiCompatibleRegistryEntry (base https://api.lyceum.technology/openai/v1, live /models discovery, passthrough). Generic DefaultExecutor handles chat/embeddings — no custom executor or translator needed. - apikey metadata card + AGGREGATOR_PROVIDER_IDS + PROVIDER_ENDPOINTS. - lyceumQuotaFetcher: reads the credit balance from GET /api/v2/external/billing/credits (pay-per-use), surfaced as a "credits" window in Dashboard > Limits and quota-aware preflight. Registered via quotaTrackersBatch (keeps chat.ts frozen). - usage leaf + dispatch + fetcher/supported/apikey-limits lists + label. - tests: registry-shape (wave1-c) + 16-case quota fetcher/usage suite. - docs: changelog fragment, regenerated PROVIDER_REFERENCE, provider count 354->355 across README/AGENTS/llm.txt (+42 i18n mirrors)/SVGs, file-size baseline bump for the gateways.ts catalog entry. * docs(providers): regenerate provider reference for Lyceum (359→360) Rebased onto release/v3.8.51 after #12462 (LLM Gateway DevPass quota) merged and the base advanced. The live registry now totals 360 providers, so the auto-generated reference and the diagram/marketing counts are regenerated against the new base: - npm run gen:provider-reference → docs/reference/PROVIDER_REFERENCE.md - 359→360 in README.md, AGENTS.md, llm.txt (provider count), package.json description and the diagram SVGs (readme-hero, promise-pillars, comparison-table, cli-terminal, tier-flow-dark/light) - node scripts/i18n/sync-llm-mirrors.mjs → 65 locale mirrors Addresses the maintainer review asking to rerun gen-provider-reference.ts and push the doc files missing from the original diff, which left Docs Gates red. * test(providers): refresh count assertions + golden snapshot for the new Lyceum provider
* feat(sse): add LLM Gateway DevPass quota tracking Surface the LLM Gateway DevPass allowance (GET /v1/key) in OmniRoute's quota telemetry, mirroring the OpenRouter API-key fetcher pattern. - llmgatewayQuotaFetcher.ts: fetch + parse the DevPass /v1/key response (decimal-string USD values), exposing two windows — monthly plan credits and the 7-day premium-model window — with a 45s TTL cache. Pay-as-you-go keys (devPlan "none") and 401/403 fail open (no quota). - Register in chat.ts before registerGenericQuotaFetchers + register the named windows for the dashboard cutoff modal. - usage/llmgateway.ts leaf + usage.ts dispatch case so the Limits page renders the monthly + weekly premium rows. - Add "llmgateway" to USAGE_FETCHER_PROVIDERS, USAGE_SUPPORTED_PROVIDERS, PROVIDER_LIMITS_APIKEY_PROVIDERS, and the dashboard label/order map. - tests: 21 cases covering the parser, auth fail-open, cache TTL, window exhaustion, preflight proceed/block, registration, and the usage leaf. * docs(sse): add changelog fragment + codebase-doc entry for llmgateway quota * refactor(sse): register llmgateway quota via quotaTrackersBatch Move the LLM Gateway fetcher registration out of chat.ts (a frozen file-size-baseline chokepoint) into quotaTrackersBatch.ts, the dedicated side-effect module that exists precisely so new fetchers don't grow chat.ts. The batch import runs at module load, before registerGenericQuotaFetchers(), so the bespoke fetcher still wins over the generic path. Fixes the file-size gate (chat.ts must not grow). --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Summary
Adds DevPass quota tracking for the
llmgatewayprovider, so the LLM Gateway subscription allowance shows up in OmniRoute's quota telemetry (Dashboard → Limits, quota-aware preflight/monitor) alongside OpenRouter, OpenCode, etc.LLM Gateway exposes a single documented endpoint —
GET /v1/key— authenticated with the connection's own gateway API key (llmgtwy_…). It returns the DevPass allowance as decimal-string USD values. This PR mirrors the existingopenrouterQuotaFetcher.tsAPI-key pattern.What it surfaces
Two independent DevPass windows:
devPlanCreditsUsed/devPlanCreditsLimit/devPlanCreditsRemaining(the recurring DevPass allowance).devPlanPremiumWeeklyLimit/devPlanPremiumCreditsUsed/devPlanPremiumWeekResetsAt(7-day rolling; resets to full allowance when the window expires).Edge cases handled (per the upstream docs):
devPlan: "none") return no quota (treated as unlimited).401(invalid/inactive key) and403(publishable/session key) fail open — quota tracking never blocks routing.Changes
open-sse/services/llmgatewayQuotaFetcher.ts(new) — fetch + parse/v1/key, build aQuotaInfowith monthly + weekly premium windows, cache, andregisterLlmgatewayQuotaFetcher().open-sse/services/usage/llmgateway.ts(new) — dashboard usage leaf returning the twoUsageQuotarows.src/sse/handlers/chat.ts— register the bespoke fetcher beforeregisterGenericQuotaFetchers()(so it wins over the generic path) + its named windows.open-sse/services/usage.ts+usage/fetcherProviders.ts— dispatch case + registration list entry.open-sse/services/usage/supportedProviders.ts+src/lib/usage/providerLimits.ts— addllmgatewaytoUSAGE_SUPPORTED_PROVIDERSandPROVIDER_LIMITS_APIKEY_PROVIDERS(otherwise/api/usage/[connectionId]rejects the connection).ProviderLimits/constants.ts— dashboard label + order.tests/unit/llmgateway-quota-fetcher.test.ts(new) — 21 cases.Testing
tests/unit/llmgateway-quota-fetcher.test.ts— 21/21 pass (parser incl. decimal strings + expired premium window; 401/403/network fail-open; pay-as-you-go → null; cache TTL; monthly & weekly exhaustion →limitReached; registration wires preflight/monitor/windows; preflight proceed vs block; usage leaf rows).generic-quota-fetcher(10/10),provider-limits-ui(25/25),free-tier-providers-wave1-c(2/2) still green — bespoke registration doesn't disturb the generic ordering.npm run typecheck:coreclean.check:provider-consistencyOK.