feat: route every chat alias to a free OpenRouter model, and halve hive-default and hive-auto - #1099
Merged
Merged
Conversation
…half price Owner directive, 2026-08-23: route both aliases to OpenRouter free models and charge 50 percent of the current price. Prices halve exactly, with no remainder and no rounding decision: hive-default goes from 10500 in and 42000 out to 5250 and 21000, hive-auto from 21000 and 84000 to 10500 and 42000, both in credits per million tokens. The old figures come from 20260822_02 step 7, the statement that set them. Both aliases route to dots-studio/dots-3-note-preview:free, which is the only zero-priced model on OpenRouter today that is simultaneously full parity on tools, tool_choice, response_format and structured outputs, live verified working, and served by a provider that does not train on prompts. The route carries provider.data_collection deny as a fail closed guard. No Go change on the serving path. precedence.go bills prompt and completion tokens from the two alias price columns in one place, so halving those columns halves every charge exactly, including the byte estimated fail closed shape, and never to zero. Which token classes are billed does not move. route-free-auto carries supports_batch, supports_image_generation and supports_image_edit forward from route-groq-auto, which is their sole carrier in the catalog. Without that, disabling it would leave /v1/batches, /v1/images/generations and /v1/images/edits with zero eligible routes for every alias in the system.
…ces unchanged Second half of the same owner directive: move the Groq TEXT and chat completion models to OpenRouter free to stop the Groq free tier allowance being consumed. hive-small, hive-medium and hive-fast are repointed to dots-studio/dots-3-note-preview:free. Their customer facing prices do NOT move. The 50 percent instruction applies to hive-default and hive-auto only, so serving these three from a free upstream simply widens margin, which is the intended outcome. That is enforced rather than promised: TestGroqFreeRepointTouchesNoPrice fails if this migration writes any price column, and the migration writes to model_aliases not at all. Groq speech to text and text to speech are untouched. A survey of all 422 OpenRouter models found no model at all advertising selectable voices and no OpenAI compatible speech endpoint, so Orpheus has no replacement at any price, and the only free models accepting audio take it as chat input rather than through a transcription endpoint. GROQ_API_KEY stays required. Capability parity was probed live rather than inferred, because the free model does not list reasoning_effort, stop, seed, the penalties, top_k or logprobs. Twelve request shapes covering all of those plus json_schema structured output returned 200, so no request shape that works today fails after the repoint. Concentration risk, accepted and recorded: every chat alias except the two paid DeepSeek ones now resolves to one model at one provider on one free endpoint capped at 20 requests per minute, with no gateway fallbacks. The failure mode moves rather than disappearing, and the workaround of switching to hive-small on a different provider no longer exists. What a customer sees at the cap is issue #1089, unchanged here. This does not close issue #1088.
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Warning Review limit reachedNext included review available in 45 minutes. View limit detailsLimit details: You’ve used the included review currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (7)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This was referenced Aug 23, 2026
sakibsadmanshajib
added a commit
that referenced
this pull request
Aug 23, 2026
…ct (#1061) ## What this is A machine gate that runs after `deploy-demo-box` succeeds and asserts that the product actually works. A deploy exiting 0 is not that claim: PR #787 merged green and took chat sign-in down, and three consecutive deploys failed at `migrate` while the box silently stayed on two-week-old code. No human approval anywhere. There is no `environment:` key on any job, and there must never be one: `main` went `7cfc460ea` (added an approval gate) then `09607ff05` (removed it). The rigour comes from stronger checks, not from a person in the loop. ## What it gates on | # | Check | What would otherwise pass | | --- | --- | --- | | 1 | **Chat answers.** `GET /api/config` returns 200 with `status: true` and an `oauth.providers.oidc` entry. | The static shell serves a full 200 HTML page over a dead backend. A page title or a bare status code passes straight through an outage. | | 2 | **Sign-in completes.** A minted session must resolve at `GET /api/v1/viewer` to the **same user id** it was minted for. | A rendered login page proves nothing. Control-plane revalidates every bearer against GoTrue per request, so this cannot be faked; comparing the id rather than the status rejects a 200 carrying the wrong identity. | | 3 | **A ledger entry lands.** A real completion through a freshly minted key must produce a new `usage_charge` row with `credits_delta < 0`. | Billing failing open is invisible from every other angle: the request succeeds and the current trace would be a row that is not there. That is exactly what happened for three days in July 2026. | Plus a **preflight**, which runs first and gates too: does the Supabase configuration in front of us name a project that is alive, with keys that belong to it. And one **measurement that reports without gating**: whether every container in the stack is running the image its tag currently resolves to (issue #869). Why it reports rather than gates is in the next section. The key for check 3 is minted **through the session**, not taken from a pre-provisioned key, so key and ledger are guaranteed to share an account. It is revoked in a `finally` block, and a revocation that itself fails is now a check failure rather than a warning. ## Honest coverage: the three failure modes this was asked to catch **A deploy that reports a rebuild while the container keeps serving the old binary (#869): measured and reported, NOT gated.** The comparison is sound where it matches: 15 of 20 containers match their tag exactly, and control-plane, edge-api and open-webui moved from mismatched to matched across a deploy that recreated them, so the check tracks real state. What the image-id comparison cannot yet establish on this box is the **direction**, which is the entire difference between a stale binary and harmless drift: five containers report an image id that `docker image inspect` cannot resolve at all while their tag resolves to an image built seconds before the container was created, and a container newer than its tag's image did not miss that image. Failing a deploy on that would leave this gate red on every deploy from the day it merges, and a gate that is always red is the one nobody reads on the day it means something. So the numbers ship with the run, the evidence goes to #869, and turning it fatal is a one-line change with a comment saying exactly that. **This gate does not currently catch #869.** **A deploy that is green while shipping a dead URL (#1059 class): caught, and proved.** This is the preflight. It was, however, weaker than its own header claimed: pointed at `https://chat-hive.scubed.co`, a single-page app that answers 200 with HTML for every path it does not know, it reported the Supabase configuration sound. Measured on the live box, not imagined. It now requires the `/auth/v1/health` body to be GoTrue's own health document, and the `config` negative control reproduces the failure exactly that way, since a hostname updated in one place and left behind in another is the shape #1059 actually took. **A migration that fails while the tracker reports success: NOT caught here, and it should not be.** That check already exists upstream and closer to the event: `scripts/apply-migrations.sh` plus `scripts/probe-applied-migrations.py` decide the baseline state from the schema itself rather than from the ledger, precisely because an empty ledger has two opposite meanings, and `migrate` runs before `deploy` so a failure stops the deploy. This gate notices such a failure only if it changes chat, sign-in or billing behaviour. The gap worth closing separately is a post-deploy drift assertion (re-run the probe read-only and fail on any applied-versus-present disagreement); it is deliberately not in this pull request. ## Where each input comes from URLs and keys come from the **box**. Six workflows read `SUPABASE_*` from GitHub secrets naming a hosted project that has been deleted, and nothing went red (#1059). The box's `.env` is what the running containers actually use. Two corrections to that were needed before this could run at all: - `SUPABASE_URL` on the box is `http://caddy-supabase`, an in-network compose hostname. This job runs on the host, outside the compose network, so signing in against it dies in DNS. The auth origin is taken from `NEXT_PUBLIC_SUPABASE_URL` (the public origin the product's own browser clients use), with the anon key read from `NEXT_PUBLIC_SUPABASE_ANON_KEY` first so the origin and the key stay paired. A single-label host now fails immediately, naming why. - Cloudflare fronts both demo hostnames and answers 403 "Error 1010" to the literal `Python-urllib/3.x` User-Agent. Measured: the same `GET /api/config` returns 403 with that default and 200 with a named agent, from the same host in the same second. Both scripts send one. All three verification targets (`HIVE_CHAT_URL`, `HIVE_CONTROL_PLANE_URL`, `HIVE_EDGE_API_URL`) are required by the script from its environment with no default: a checker that silently assumes a fixed demo host can return green for a product nobody deployed. The workflow exports them explicitly, box `.env` first, with the deployment's public hostnames as a reviewed fallback whose source every run reports. The verification **identity** comes from the box's `.env`. That is not a contradiction. A stale URL is dangerous because nothing exercises it; a stale credential cannot hide, because signing in with it is the next thing this job does. ## Verification identity: configured, provisioned, proved Resolved during this session; nothing is left for the owner here. - `HIVE_VERIFY_EMAIL` and `HIVE_VERIFY_PASSWORD` exist in `/home/sakib/hive/.env` on the box (added outside this branch; presence verified by reading the file's key names only). - The account behind them was checked read-only against the database: a real GoTrue user, email confirmed, single tenant membership, own billing account. - It was not usable as first written: membership was MEMBER and balance zero, which the gate itself surfaced as a 403-shaped failure class and an upstream `insufficient_quota`. Provisioning applied two scoped writes limited to that user's own rows: membership raised to OWNER (minting an API key needs `api_keys:write`, granted to OWNER only), and one idempotent grant entry of 10,000 credits on its own billing account (`idempotency_key post-deploy-verify-seed-2026-08-23`, unique index makes reruns inert). - First fully-green verification followed immediately, on the box, all three checks passing including a real charge of `-1` credit and clean key revocation. ## Negative controls Every gating check has one, and each removes the guarded condition for real: | Value | What it removes | Expected | | --- | --- | --- | | `chat` | reads the static shell at `/` instead of the backend config | red | | `signin` | flips every bit of the bearer's first decoded signature byte, header and payload intact | red | | `ledger` | skips the completion, so nothing bills | red | | `config` | points the auth origin at another live host in this deployment | red | | `freshness` | compares against an image id nothing can be running | red | No value can make anything pass; it can only cause a failure. Selectable from `workflow_dispatch` after this merges, and from a `verify-negative:<value>` label before then. Selecting `signin` or `ledger` with no identity configured fails immediately naming that precondition rather than surfacing later as a confusing argparse error. ## Verification, run by run All against the live box unless noted. | Run / evidence | What it shows | | --- | --- | | On-box full run, 2026-08-23 ~20:42 UTC | **Full green baseline, all three checks**: chat backend config 200 with `status: true` and OIDC provider; minted session resolved at the control-plane to the same user; real completion answered 200, produced `usage_charge` `credits_delta=-1`, key revoked. | | [32665274872](https://github.com/sakibsadmanshajib/hive/actions/runs/32665274872) | `signin` control red through the assertion: corrupted decoded signature bytes, control-plane answered 401 `invalid or expired token`. | | [32665392330](https://github.com/sakibsadmanshajib/hive/actions/runs/32665392330) | `ledger` control red through the assertion: completion skipped, no new `usage_charge` within 90s, reported as billing failing open. | | [32660681673](https://github.com/sakibsadmanshajib/hive/actions/runs/32660681673) | Earlier partial green baseline: preflight passed from GoTrue v2.189.0 at the public origin, chat check passed, 20 containers compared, missing identity reported as unconfigured coverage. | | [32659829693](https://github.com/sakibsadmanshajib/hive/actions/runs/32659829693) | `chat` control red, and for the right reason: `/` answered 200 with HTML, so the check names the static shell rather than the backend. | | [32660319162](https://github.com/sakibsadmanshajib/hive/actions/runs/32660319162) | `config` control red: `GET /auth/v1/health` answered 200 from a host that is not GoTrue, and the preflight says so instead of passing. | | [32660549692](https://github.com/sakibsadmanshajib/hive/actions/runs/32660549692) | `freshness` control red through the assertion path, every container reported and the step failing on the comparison. | | [32660408662](https://github.com/sakibsadmanshajib/hive/actions/runs/32660408662) | The same control failing for the wrong reason before it was fixed. Kept because a control that fails for the wrong reason proves nothing, and this is what that looks like. | | [32634926680](https://github.com/sakibsadmanshajib/hive/actions/runs/32634926680) | The original failure this PR had to fix: no verification identity in the box `.env`. | | [32664630991](https://github.com/sakibsadmanshajib/hive/actions/runs/32664630991) | The new response-shape guard catching a real condition live: a genuine zero-row account answers `{"entries": null}`, which is the Go handler's designed empty answer; the scan now reads null as empty and any other non-list type still fails naming its type. | | [32665543894](https://github.com/sakibsadmanshajib/hive/actions/runs/32665543894), [32665768538](https://github.com/sakibsadmanshajib/hive/actions/runs/32665768538) | Upstream capacity noise, documented not hidden: after #1099 every chat alias routes to an OpenRouter free model, whose per-key daily capacity intermittently answers 429 `insufficient_quota` while chat itself stays up. The completion is retried up to `HIVE_VERIFY_COMPLETION_ATTEMPTS` (default 3) with `HIVE_VERIFY_RETRY_DELAY` (default 20s) between attempts; exhausted retries fail naming the last answer. Expect the ledger check to flap with OpenRouter free-tier capacity until the route mix changes. | ## Review - **CodeRabbit CLI**: ran twice (round 1: three findings, three fixes; round 2 on the new commits: four findings, four fixes: target URLs required rather than defaulted, JSON shapes validated before field access, byte-changing signature mutation, revocation failure fatal, secret lengths suppressed for password/service-role values, deterministic negative-control invocation). - **`ecc:code-review`**: ran as a structured seven-category pass over the three changed files at head revision. Findings folded back into the fixes above plus one consistency fix (service-role key length suppressed in the preflight output). Posted as a COMMENT review; the builder does not approve its own PR. - **Plain adversarial pass**: ran, against the live box rather than by reading. Produced the Cloudflare user-agent block, the non-GoTrue health-document gap, subset-run summary wording, notifier scope, freshness-control ordering, anon-key pairing, the entries-null discovery, and the 429 retry semantics. - **Mandatory security review on an auth and money path**: manual pass covered: no value printed anywhere (presence only for password and service-role values, masked on export), no credential crossing `$GITHUB_ENV`, a step summary or a job boundary, no password ever written or rotated (`docs/live-test-auth.md`), the minted key revoked in a `finally` with failure now fatal, fork head repositories excluded from the self-hosted runner, `persist-credentials: false`, and the real spend bounded (one completion, `max_tokens: 64`, dedicated identity, never the demo account). - **`/codex:adversarial-review`**: SKIPPED, capability gap, not a clean pass. Codex usage limits were reached earlier today (recorded on the PR timeline); the stream could not run. ## Failure handling: alert, not rollback Deliberately no automatic rollback. `migrate` runs before `deploy` and forward migrations are not reversible by redeploying code, so an automatic rollback would leave old code against a new schema, which is worse than a bad deploy. The notifier uses its own `ci-failure:post-deploy-verify` label rather than sharing `ci-failure:deploy-demo-box`, since this workflow cannot join that job's `needs` list and its ledger check can go red for reasons unrelated to a deploy. It fires **only on `workflow_run`**: issue #1062 exists because a pull request run filed it, and every negative-control run is expected to fail. It dedupes by exact title over the issues list API rather than the eventually consistent search API, and `continue-on-error` plus `|| true` mean it can never alter the verdict. ## Not a user-visible surface No screenshot: this changes a workflow and two scripts. Nothing in the product's UI moves, so the visual-proof rule does not apply. The evidence is the run table above. ## Test plan - [x] Green baseline: the checks that can run pass against the running box (partial CI baseline, then a full three-check green run on the box once the identity was provisioned). - [x] `negative_control=chat` goes red. - [x] `negative_control=config` goes red, restore goes green. - [x] `negative_control=freshness` goes red through the assertion. - [x] `negative_control=signin` goes red through the assertion (run 32665274872). - [x] `negative_control=ledger` goes red through the assertion (run 32665392330). - [x] Preflight fails legibly on a Supabase configuration that answers HTTP but is not our project. - [x] The notifier cannot fire from a pull request run. - [x] CodeQL alert 27 cleared, thread resolved. - [ ] One more fully-green CI run once OpenRouter free-route capacity allows a completion through; the on-box full run already proves the path end to end. ## Buglog entry To be appended to `.wolf/buglog.jsonl` on `main` in a buglog-only pull request after this merges, per issue #873. ```json {"date":"2026-08-23","title":"Cloudflare answered 403 Error 1010 to the default Python urllib user agent, so every HTTP check against the demo hostnames failed","error_message":"HTTP 403 {\"title\":\"Error 1010: Access denied\",\"detail\":\"The site owner has blocked access based on your browser's signature\"}","root_cause":"urllib sends User-Agent: Python-urllib/3.x by default, and the Cloudflare bot rules in front of chat-hive.scubed.co and console-hive.scubed.co block that signature. curl from the same host in the same second returned 200, which is why the failure read as a dead backend rather than a blocked client.","fix":"scripts/post-deploy-verify.py and scripts/preflight-supabase-config.py set an explicit named User-Agent on every request.","tags":["cloudflare","http","verification","false-negative","demo-box"]} {"date":"2026-08-23","title":"The Supabase preflight reported a wrong host as a healthy project because it only checked the status code","error_message":"PREFLIGHT PASSED against SUPABASE_URL=https://chat-hive.scubed.co","root_cause":"The liveness check asserted GET /auth/v1/health returned 200 and nothing about the body, while chat-hive.scubed.co is a single-page app that answers 200 with HTML for every unknown path.","fix":"The preflight requires the health body to be GoTrue's own document and fails naming what answered instead.","tags":["supabase","preflight","false-positive","issue-1059"]} {"date":"2026-08-23","title":"post-deploy-verify filed an issue claiming the demo box was broken from a pull request run","error_message":"issue #1062 post-deploy-verify is failing against the demo box, verify=failure, trigger=pull_request","root_cause":"The report-failure job was gated on failure() alone, so a label-gated pull request run of the workflow, including every negative-control run that is expected to fail, filed an outage issue for the deployment.","fix":"The notifier requires github.event_name == 'workflow_run', and a missing verification identity files under its own title.","tags":["github-actions","notifier","false-alarm","issue-1062"]} {"date":"2026-08-23","title":"The verification could not sign in at all because the box's SUPABASE_URL is an in-network compose hostname","error_message":"SUPABASE_URL=http://caddy-supabase does not resolve from the runner host","root_cause":"Services inside the compose network reach auth at http://caddy-supabase; this job runs on the host outside that network. The public auth origin is NEXT_PUBLIC_SUPABASE_URL, a different value on this box.","fix":"The workflow reads NEXT_PUBLIC_SUPABASE_URL first, pairs it with NEXT_PUBLIC_SUPABASE_ANON_KEY, and fails immediately when the resolved host has no dot in it.","tags":["supabase","demo-box","networking","verification"]} {"date":"2026-08-23","title":"The verifier crashed with AttributeError instead of a named failure when a 200 carried the wrong JSON shape","error_message":"AttributeError on payload.get / viewer.get / entries iteration","root_cause":"Decoded JSON was field-accessed without type validation, and main() catches CheckFailed only, so a malformed 200 skipped the per-check summary entirely.","fix":"Sign-in, viewer and ledger payloads are shape-checked before access and raise the named CheckFailed.","tags":["python","verification","error-handling"]} {"date":"2026-08-23","title":"A real zero-row ledger answered {\"entries\": null} and the new shape guard rejected a correct API answer","error_message":"ledger read answered 200 but carried entries as NoneType, expected a list","root_cause":"The control-plane handler builds the list with var entries []LedgerEntry, so an empty account marshals as null, not []; the guard written for malformed bodies treated the designed empty answer as malformed. Found live on the first full run against the box.","fix":"Null entries scan as the empty page; any other non-list type still fails naming its type.","tags":["go","json","verification","null-semantics"]} {"date":"2026-08-23","title":"Ledger check flapped on upstream 429 after the chat aliases moved to an OpenRouter free route","error_message":"the completion answered 429 insufficient_quota on attempts 1..3 while an identical completion minutes earlier returned 200 and charged normally","root_cause":"#1099 repointed every chat alias at an OpenRouter free model whose per-key daily capacity intermittently refuses service; the verifier's claim is that a served request bills, not that an upstream always serves.","fix":"Bounded retries (HIVE_VERIFY_COMPLETION_ATTEMPTS=3, HIVE_VERIFY_RETRY_DELAY=20s) treat 429/5xx as retryable; exhausted retries fail naming the last answer. Documented as expected flap until the route mix changes.","tags":["openrouter","rate-limit","verification","flaky-gate"]} ``` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added automated post-deployment verification after successful deployments, with support for manual runs and eligible pull requests. - Verifies chat availability, sign-in, billing completion, container freshness, and deployment configuration. - Supports targeted checks, negative testing, retries, polling, and safe credential handling. - Added Supabase configuration preflight checks for environment settings, tokens, URLs, and service health. - Reports skipped checks when credentials are unavailable without incorrectly failing verification. - **Bug Fixes** - Added deduplicated issue notifications for verification failures or missing coverage without changing verification results. <!-- end of auto-generated comment: release notes by coderabbit.ai --> ## Billing test evidence (CodeRabbit follow-up) - `go test ./apps/edge-api/... -count=1 -short` in the toolchain container: all packages pass, including the reservation guard tests covering the 429 mapping (`reservation_guard`) and the strict-mode 409 to 429 translation. - `go test ./apps/control-plane/... -count=1 -short`: all packages pass, including the reservations policy tests for the flat hold and PolicyError path (`service.go` strict mode). - Live evidence: run 32669144332 rerun completed success at 2026-08-23T23:10:22Z with the full three-check green (chat, signin, ledger), after the verify workspace balance was topped up past the flat 10000 credit hold. The named preflight added in stacked PR #1105 makes this failure mode self-diagnosing going forward.
sakibsadmanshajib
added a commit
that referenced
this pull request
Aug 23, 2026
…1105) Stacked on #1061 (targets its branch, since scripts/post-deploy-verify.py exists only there). Not for merge until #1061 lands, then rebase or retarget to main. ## What happened live (2026-08-23) The post-deploy-verify ledger check failed every run with: ``` HTTP 429 {"error":{"message":"You exceeded your current quota, please check your plan and billing details.","type":"insufficient_quota","code":"insufficient_quota"}} ``` The wording reads like an upstream provider quota error. It is not. Every upstream layer tested clean: direct OpenRouter 200, LiteLLM route-deepseek-v4-flash 200, catalog alias mapping correct, zero requests reaching litellm during the failures. ## Root cause 1. The verify workspace was seeded with exactly 10000 credits (`post-deploy-verify-seed-2026-08-23` grant). 2. The first passing run consumed 1 credit, leaving 9999. 3. Every chat completion reserves a flat **10000 credit hold** before dispatch (`apps/edge-api/internal/inference/chat_completions.go` passes that endpoint default to `ReservationCredits`, which falls back to it because `reservation_estimate_credits` is null on every fixed-price alias). 4. Control-plane refuses the reservation with a strict-mode PolicyError 409 (`reservation exceeds available credits`, `apps/control-plane/internal/accounting/service.go`). 5. Edge maps that 409 to HTTP 429 `insufficient_quota` with the OpenAI-style message text (`apps/edge-api/internal/inference/reservation_guard.go`). No usage_events row is written because nothing was ever dispatched. Refused requests consume nothing, so the balance never recovers and every rerun fails identically. The timing correlation with #1099's deploy was coincidence. ## What this PR changes A named preflight in `check_ledger` before any spend: read `GET /api/v1/accounts/current/credits/balance` and fail with the real reason plus remediation when available credits are below the flat hold. A non-200 read warns and proceeds so the completion attempt still speaks for itself; the negative control skips both spend and preflight. ## Evidence - Reproduced through the public edge with a key on the verify workspace at balance 9999: exact 429 body above. - Fixed operationally by a +10000 grant ledger row (`post-deploy-verify-topup-2026-08-23`); same call then answered 200. - Rerun of post-deploy-verify run 32669144332 completed success after the top-up. ## Buglog entry {"error_message":"post-deploy-verify ledger check fails with HTTP 429 insufficient_quota from /v1/chat/completions","root_cause":"verify workspace balance 9999 credits fell below the flat 10000 credit chat reservation hold after the first passing run consumed one credit; control-plane answers 409 policy rejection and edge maps it to a 429 whose message mimics an upstream provider quota error","fix":"operational +10000 credit grant to the verify workspace ledger; PR adds a named balance preflight so the gate reports this condition instead of an upstream-looking 429","tags":"billing,reservation,fail-closed,post-deploy-verify,misleading-error"} ## Test plan - [x] py_compile passes - [x] Live reproduction of the 429 at balance 9999 and 200 after top-up - [x] Gate rerun green against the topped-up box - [ ] Negative path: balance below hold produces the new named CheckFailed (exercisable by draining the workspace again)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Owner directive, 2026-08-23: route both
hive-defaultandhive-autoto OpenRouter free models, and charge 50 percent of what is being charged now.One migration plus one LiteLLM config edit. No Go change on the serving path.
First, a correction to the premise this task was handed with
The brief said
hive-autobills at actual upstream cost after PR #1012, and therefore had to be converted back to a fixed price with the "today" figure recovered from ledger rows.That is not what #1012 did. It created a new alias,
openrouter-auto, on a new routeroute-openrouter-auto-beta, inpricing_mode = 'upstream_actual'. Its own migration says so out loud: "litellm_model_name is deliberately NOT 'route-openrouter-auto': that name is the retired route id of the pre-existing hive-auto alias, which resolves to a completely different model."hive-autowas never touched by it. Both aliases are plainfixedrows and have been since the 2026-08-22 restructure, so there is no pricing-mode conversion here and no need to reconstruct a price from ledger magnitudes. The old figures come from the statement that set them.Before and after
Source for every old rate:
supabase/migrations/20260822_02_catalog_alias_restructure.sqlstep 7.Every figure halves with no remainder, so no rounding rule had to be invented and no float can enter. Both aliases stay
pricing_mode = 'fixed'andprice_unit = 'tokens'.Why the cache columns are listed at 0 and not omitted: they are the only other price columns on the row,
precedence.gonever reads them, and both were already zeroed by 20260822_02 when the aliases moved to Groq routes declaring no cache support. Half of zero is zero, so they are asserted rather than changed. Naming them is the point: "every unit type the alias bills today" is answered exhaustively rather than by omission.Why there is no Go change
apps/edge-api/internal/metering/precedence.gois the single implementation of the charge arithmetic (D-031). It builds exactly twoUnitChargevalues per request, prompt tokens atInputPriceCreditsand completion tokens atOutputPriceCredits, sums them, divides once by a million and rounds half up. So halving those two columns halves every token charge exactly, for every request shape, with nothing else moving:ReservationCreditsreturns the flat per-endpoint default (10000 for chat); it only raises a hold above that for anupstream_actualalias. The hold is an authorization, released in full at settlement./v1/images/*reserves a hardcoded 5000 credits andinternal/audiocharges flat literals, both alias-independent (D-033, issue Audio billing ignores the catalog and charges flat hardcoded credits #627). See the residuals.The money proof: same request shapes, before and after
Produced by calling the production settlement function
inference.CreditsForTokensdirectly, once with the old rates and once with the new ones, in the repo's own toolchain container. Not a reimplementation: that is the function both the streaming and the sync settlement paths call.Two figures per shape. Exact is quantity times rate, summed, before the single division and round-half-up, which is where "exactly half" is provable. Settled is the whole-credit charge the ledger records, and both sides round independently.
Exactly half on every row, in integer rational arithmetic. Not "roughly half", not a fraction of a percent, not unchanged, and not zero.
Stated rather than smoothed over: the settled whole-credit charge can sit one credit above exactly half, because the old and new charges each round half up independently from exact rationals in a 2 to 1 ratio (29.4 rounds to 29 while 14.7 rounds to 15). The deviation is bounded at one credit, which is 0.00001 USD at the repo's 100000-credits-per-USD constant. The 1-credit floor in
CreditsForTokensis the other bounded exception: it is what keeps a sub-credit request off zero, it predates this change, and it is why the two floor rows read 1 and 1. Both are visible in the table rather than hidden behind a "50 percent" claim the integers do not literally satisfy.The replay enforced its own claims: it failed the run if any exact ratio was not exactly one half, if any new charge was zero where the old was positive, or if any new charge exceeded half by more than one credit.
Full capture, including the commands:
docs/proof/free-route-aliases-half-price-2026-08-23/README.md.The test that fails on the old rate
apps/control-plane/internal/routing/free_alias_pricing_test.go, offline, reusing the SQL parser already in that package. Ten guards, all positional: the migration declares its arithmetic in a-- HALVE| alias | field | old | newtable, the test checks that table is arithmetically half, that its OLD column matches the rates actually in force (pinned here so a later edit to 20260822_02 cannot move the baseline), and then that the value eachUPDATEassigns equals half. A correct comment above a wrongUPDATEis caught.Mutation tested rather than asserted to work. Six mutations, six kills, control green:
TestFreeAliasPricesAreExactlyHalfTheOldRatesTestFreeAliasPricesAreExactlyHalfTheOldRatessupports_batchand both image flagsTestFreeRouteAutoCarriesTheSoleCapabilityFlagsForward:freesuffixTestFreeAliasRoutesTargetAFreeOpenRouterModeltools_supportedTestFreeRoutesKeepTheCapabilitiesTheirAliasesServeTodayTestRetiredGroqRoutesAreDisabledAndRepointedM1 is the guard the brief asked for: red on the old rate, green on the new one.
The existing
TestCatalogAliasPricesMatchProviderRatescannot cover these two prices, and that is structural rather than an oversight. It derives every credit figure asusd_per_million * 1.4 * 100000(D-032), and itsparseRaterefuses a zero rate outright: "that is a mispricing, not a rate". A free upstream costs zero, so the margin formula yields zero, and zero is refused bySelectRouteas unpriceable. These two prices are owner-set, not cost-derived, and the halving relation is what replaces the formula as the checkable invariant.Routing: which free model, chosen from live data
https://openrouter.ai/api/v1/models, fetched 2026-08-23: 422 models, 22 priced at zero on both prompt and completion. A literalopenrouter/freeid does exist; it is a router, not a model.The capability bar comes from what these aliases serve today. Both current routes declare
tools_supported = true, which is the column PR #206 routestools,tool_choiceandresponse_formaton, plussupports_streamingandsupports_reasoning. Of the 22 free models, five support all of tools, tool_choice, response_format and structured outputs. Joined tohttps://openrouter.ai/api/frontend/v1/all-providers:A false green I shipped and then caught, recorded so the next person does not repeat it: the endpoints API reports NVIDIA's provider as
Nvidiawhile the provider directory keys it under displayNameNVIDIA. Joining on displayName alone misses the record and reports every NVIDIA free endpoint as no-training and zero-retention, the exact opposite of the truth. Join on both fields.Live probes with the project's real key, and this is where the paper ranking fell apart:
z-ai/glm-5.2:free, the only zero-retention candidate and the strongest model on every published axis: 429 on four of four attempts,provider_error_code: upstream_429,limit_source: upstream_provider_shared_pool. That is Decart's shared pool, not our account's limit, and buying credits cannot raise it. Rejected on live evidence rather than on paper.openrouter/freewith no provider preference: five of five succeeded, but four landed on NVIDIA endpoints and two of five landed onnvidia/nemotron-3.5-content-safety:free, a moderation classifier, which answered a plain chat prompt withUser Safety: safeand then with an empty string. Unfit for a chat alias on output quality alone, before the training question. This is also the exact shape issue hive-default and hive-auto route through OpenRouter's random free-model router, not the models their catalog prices are derived from #689 called a bug.openrouter/freewithprovider: {data_collection: "deny"}: five of five, only Cohere and AtlasCloud, tool calls on five of five. The deny filter empirically excludes every NVIDIA endpoint and the classifier.provider: {zdr: true}, the stricter form: 404, "No endpoints found matching your data policy (Zero data retention)". There is no zero-data-retention free endpoint reachable at all today.dots-studio/dots-3-note-preview:free: 200 on a sync completion, on a tools request (a realtool_callswith correct arguments andfinish_reason: tool_calls), onresponse_format: {"type":"json_object"}(valid JSON back) and on a streamed request.cost: 0on every response.Chosen:
dots-studio/dots-3-note-preview:freefor both aliases, pinned, withextra_body.provider.data_collection: denyandallow_fallbacks: false. It is the only free model that is simultaneously full parity, live verified working, and served by a provider that does not train on prompts. Thedenypreference is kept even though the model resolves to one provider today: if AtlasCloud's policy changes or a second provider appears, the request fails instead of quietly moving customer prompts to a provider that stores them.Capability parity verdict
tools_supportedtext+image->textTwo notes on that table.
provider_capabilitieshas no vision column, so the free model's image-input support is an undeclared property of the upstream rather than a new product claim; the 2026-08-22 note that no customer-reachable vision path exists is now inaccurate at the upstream level and the config comment says so. And the three media flags are a documented status-quo fiction inherited through two migrations: neither gpt-4.1-mini, nor gpt-oss-120b, nor this model generates images. They are carried becauseroute-groq-autois their sole carrier in the catalog,SelectRoutehard-filters on each flag, andbatchstoresendsNeedBatch = truefor every batch, so disabling it without handing them on would leave/v1/batches,/v1/images/generationsand/v1/images/editswith zero eligible routes for every alias in the system. M3 above is the guard.Under-claiming is not a safe default here: both aliases are
pinnedto exactly one route, somatchesRequestedCapabilitiesdrops the only candidate,SelectRoutereturnsErrRouteNotEligibleandwriteRoutingErrormaps it to 422. On a pinned alias an under-claim is a failed request, not a withheld feature.Rate limits, and what a user sees at the cap
Documented (
https://openrouter.ai/docs/api_reference/limits.md, whose MDX constants resolve to real numbers): free variants are capped at 20 requests per minute, and at 50 per day below 10 dollars of lifetime credit purchases or 1000 per day at or above it.GET /api/v1/keyreportsis_free_tier: falsefor this account and project memory records a 10 dollar purchase, so 1000 per day is the expected tier. The key endpoint does not expose the purchased total, so that is an expectation, not a measurement; its ownrate_limitfield is documented as deprecated and returnsrequests: -1. Separately and more sharply, the Decart 429 shows a per-provider shared pool can refuse everything regardless of our standing.What a customer sees at the cap today: a roughly 60 second wait and then a 502 whose message is
context canceled. That is issue #1089, filed today and unchanged by this work.num_retries: 3withrequest_timeout: 45retries a rate-limited deployment three times, the SDK suites time out at 60 seconds, and the provider's own 429 with its retry hint never leaves the LiteLLM container log.Not fixed here, and the reason is containment rather than appetite. The candidate fix (
router_settings.retry_policy.RateLimitErrorRetries: 0) belongs in thelitellm_settingsblock, which the config sync preserves verbatim on a live volume. A file edit there is inert on the box and cannot be verified from a developer machine with no SSH to it. #1089 reaches the same conclusion about itself and asks for a real reproduction. This change does make it materially more likely, since 20 requests per minute is much tighter than Groq's ceiling, so its priority moves from latent to likely.Mitigation that does exist:
hive-smallandhive-mediumstay on Groq and are now the only customer-reachable Groq chat routes besides the deprecatedhive-fast. A rate-limited customer has a working alternative, but has to select it; nothing routes them there, because the gateway chat fallbacks were removed deliberately in 20260822_02 and re-adding one is the same inert-on-a-live-volume problem.Logging and training terms
training: false,retainsPrompts: true, no retention period published.provider.data_collection: denyas a fail-closed guard.zdr: true404. The best available free posture is a provider that does not train but does retain for an unpublished period. The owner directed the move with this on the record.Why new route ids rather than repointing the two Groq rows
The same reason 20260822_02 gave, and it applies in this direction too. The LiteLLM config sync merges field by field: the database owns only
model,api_baseandapi_key, and every other key already on the entry survives so that hand-tuning sticks (mergeParams, issue #707). Retiring the route id makes the merge drop the whole stale entry, because a knownroute_idthat is no longer active is deleted rather than updated. The rows are disabled, not deleted, so the change is reversible andSelectRoutefilters them out.price_classstaysstandard, identical to the routes being replaced.budgetwould arguably describe a free upstream better, butprice_classfeedsallow_price_class_widening, and keeping the same value means this repoint cannot change selection behaviour through a second mechanism.The
api_keyfollows automatically fromproviders.api_key_envfor the row's provider slug, soopenrouterresolves toOPENROUTER_API_KEYwithout the migration naming a secret. The doubledopenrouter/prefix onprovider_modelis correct: LiteLLM strips the leading one as its provider selector, exactly asopenrouter/~deepseek/...andopenrouter/openrouter/auto-betaalready do. The trailing:freeis load-bearing, not cosmetic: dropping it selects a paid endpoint of the same model, so the alias would charge a halved price against a real out-of-pocket cost and nothing else in the tree would notice. M4 is the guard.Verification
go test ./apps/control-plane/... -count=1 -short: clean, no FAIL, no panic.go test ./apps/edge-api/... -count=1 -short: clean, no FAIL, no panic.npm run lint:litellm-config: PASS.npm run lint:litellm-routing: PASS.npm run lint:proof-tokens: ok, 137 files scanned.deploy/litellm/config.yamlparses and resolves 14 model entries, with both new routes carrying the intendedextra_body.Not proved, and it matters: that the running gateway serves the new route. The config is volume-seeded and the live change arrives through
POST /internal/litellm/syncreadingprovider_routes. There is no SSH to the demo box from here and CI is the only remote hands, so the on-box confirmation belongs to the deploy run.deploy-demo-box.yml's "Assert model catalog prices agree with the model LiteLLM will call" step is the check that catches a stale volume, and its predicate (pricing_mode = 'upstream_actual' OR input_price_credits > 0) covers both of these rows. No ledger row at the new price exists yet either; that needs a served request after the migration applies.No screenshot, because this change alters no UI surface. The console catalog table renders whatever price the API returns, and the two figures it will show are the ones proved above.
Residual questions for the owner
Both are stated rather than silently resolved, per the brief.
hive-autonow costs exactly twicehive-defaultfor the identical model. Both resolve to the same free upstream, so the surviving 2x gap buys the customer nothing. The directive fixes each alias at 50 percent of its own old price, which is what shipped; equalising them is a separate decision. Options: (a) leave as is, honest to the directive, indefensible to a customer who compares them; (b) equalisehive-autoto 5250 and 21000, a further reduction that cannot overcharge anyone; (c) givehive-autoa distinct larger free model. Recommendation: (b). Option (c) is blocked today: every other full-parity free model is served by a provider that trains on prompts, and the one zero-retention candidate returns 429 on every request.The image path is not halved.
/v1/images/generationsand/v1/images/editsreserve and settle a hardcoded 5000 credits that never reads the alias price, so an image request throughhive-autois billed the same as before. It is also not actually servable, since the upstream is a text model and those capability flags have been a documented fiction since 20260414_01. Options: (a) halve the literal to 2500, which halves it for every other alias too and so exceeds this directive's scope; (b) leave it and close the fiction under issue Audio billing ignores the catalog and charges flat hardcoded credits #627; (c) drop the flags, which deletes three endpoints catalog-wide. Recommendation: (b).Buglog entry
{"id":"bug-2026-08-23-openrouter-provider-name-vs-displayname-join","date":"2026-08-23","title":"Joining OpenRouter endpoint provider_name against all-providers displayName silently reports a training provider as zero-retention","error_message":"A capability-and-data-policy scan of OpenRouter's free models reported every NVIDIA free endpoint as training:false, retainsPrompts:false, when NVIDIA's published policy is training:true, retainsPrompts:true","root_cause":"GET /api/v1/models/{slug}/endpoints reports the provider in its `name` form (`Nvidia`) while GET /api/frontend/v1/all-providers keys the record by `displayName` (`NVIDIA`). 18 of 82 providers have name != displayName. A dictionary keyed on displayName therefore misses the lookup entirely, and a default of an empty policy object reads as no-training and zero-retention rather than as unknown, so the failure presents as the safest possible answer instead of an error. The scan looked authoritative and was inverted on exactly the axis a data-sovereignty product cares about.","fix":"Key the provider lookup on both `name` and `displayName`, and treat a missing policy as UNKNOWN rather than as permissive. Re-ran the scan: only two of five full-capability free models are served by non-training providers, not four. Also verified the conclusion independently against live behaviour: with provider.data_collection deny set, five of five requests routed away from every NVIDIA endpoint, which agrees with the corrected join and contradicts the original one.","tags":["openrouter","data-policy","sovereignty","false-green","join-key-mismatch","provider-catalog"]}Part two: the remaining Groq text routes, at prices deliberately unchanged
Second owner directive the same day, folded into this PR because it touches the same catalog rows and the same
deploy/litellm/config.yaml, so a separate PR would conflict by construction.Move the Groq TEXT and chat-completion models to OpenRouter free as well, to stop the Groq free-tier allowance being drained.
Shipped as a second migration,
20260823_21_groq_text_routes_to_openrouter_free.sql, deliberately separate from the repricing one. Splitting them is what makes "no price moves here" a structural property of a file rather than a claim in its header.Prices for these three do NOT change, and that is the intended outcome
The owner's 50 percent instruction applies to
hive-defaultandhive-autoonly. Serving a same-priced alias from a free upstream widens margin, and that is intended. Stated explicitly so no later reader mistakes it for an oversight, and enforced two ways rather than promised:TestGroqFreeRepointTouchesNoPricefails if that migration assigns any ofinput_price_credits,output_price_credits, either cache column,pricing_modeorprice_unit. The migration does not writemodel_aliasesat all, so the guard holds in its strongest form.hive-fast's cache columns stay at 1 and 4, the stale OpenRouter-era values 20260822_02 examined and deliberately left alone so a deprecated alias would not be repriced on any axis. Not touched here either, for the same reason.Audio is out of scope, and the survey behind that
route-groq-stt(whisper-large-v3) androute-groq-tts(Orpheus) are untouched, andTestGroqFreeRepointLeavesAudioOnGroqfails if that migration so much as names them in an executable statement.GROQ_API_KEYstays required; what Groq no longer serves is chat."OpenRouter has no audio" would have been an overstatement, so here is the actual picture across all 422 models:
supported_voices. The only entries with audio in their output modality aregoogle/lyria-3-pro-previewandgoogle/lyria-3-clip-preview(free, MUSIC generation, no voice selection) andopenai/gpt-audioandopenai/gpt-audio-mini(PAID, speech-to-speech chat). None is an OpenAI-compatible/v1/audio/speechendpoint, which is whatinternal/audiospeaks. Orpheus has no replacement at any price.thinkingmachines/inkling:free,thinkingmachines/inkling-small:free,nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free). They are not a transcription endpoint: they are chat-completions models, whileroute-groq-sttis a LiteLLMmode: audio_transcriptionroute that edge-api's audio handler forwards multipart audio to. Using one would be a new integration, not a repoint. All three are also served by providers whose published policy is training on prompts, a poor destination for dictated speech specifically.Reported, not acted on, exactly as the directive asked. Bengali voice dictation (PR #1079) keeps working.
Capability parity, probed live because the parameter lists differ
This one genuinely needed checking rather than asserting.
dots-3-note-previewliststools,tool_choice,response_format,structured_outputs,reasoning,include_reasoning,max_tokens,temperatureandtop_p, and does not listreasoning_effort,stop,frequency_penalty,presence_penalty,seed,top_korlogprobs. The Groq gpt-oss models it replaces accept several of those. So: does an unlisted parameter fail, or is it ignored?Twelve request shapes, live:
All twelve returned 200 with a correct answer. There is no request shape that works today and fails after the repoint, which is the regression this check exists to rule out. The
require_parameterscase is the interesting one: OpenRouter considersreasoning_effortsatisfied by this endpoint even under strict parameter enforcement.Two behavioural differences that are not failures, recorded so nobody files them as new bugs:
reasoning_effortis accepted and never rejected, but does not reliably modulate effort:lowproduced more reasoning tokens thanhighin the same run. Requests keep working; the knob stops being meaningful.n=2returns one choice. That is OpenRouter's existing behaviour for a provider that does not implementn, identical on the routes being replaced.Capability FLAGS are carried across per route rather than uniformly.
route-free-smallandroute-free-mediummirror their Groq originals includingsupports_reasoning = true;route-free-fastkeepssupports_reasoning = false, which is status-quo preservation of an under-claimroute-groq-fasthas carried since its original 20260331_02 seed and that 20260822_02 examined and deliberately left alone. Widening it would be safe, but a routing migration is not where a deprecated alias should quietly gain a feature.None of these three is a sole carrier of a media flag, so disabling them removes no endpoint.
TestRepointedGroqTextRoutesKeepTheirCapabilitiesasserts both directions: the flags they must have, and that they claim none of the three media flags that belong onroute-free-auto.Concentration risk, stated rather than buried
After this change every customer-reachable chat alias except the two paid DeepSeek ones resolves to ONE model at ONE provider on ONE free endpoint. Five aliases:
hive-default,hive-auto,hive-small,hive-medium,hive-fast. There are no gateway fallbacks, by deliberate design (20260822_02 removed them because a fallback answers from a model the alias was not priced against).So the failure mode moves rather than disappearing. OpenRouter documents free variants at 20 requests per minute, and 50 or 1000 per day depending on lifetime credit purchases; that per-minute ceiling is tighter than the Groq daily allowance this was meant to escape. Separately, the Decart 429 recorded in part one shows an upstream provider shared pool can refuse everything regardless of our account standing, and buying credits cannot raise it.
It also removes the mitigation part one could still point at. An hour ago a rate-limited customer could select
hive-smalland land on a different provider. There is no such alias now.What a demo user sees at the cap is unchanged: a roughly 60 second wait and then a 502 whose message is
context canceled, never the provider's own 429 with its retry hint. That is issue #1089. Not fixed here for the containment reason given in part one: the candidate fix lives in thelitellm_settingsblock the config sync preserves verbatim, so a file edit is inert on a live box and cannot be verified from a developer machine with no SSH to it. This change raises that issue's priority from latent to likely.This does not close issue #1088. That is CI consuming the live demo's provider allowance, a different cause with a different owner, and there is deliberately no
Fixes #1088line anywhere in this PR.Part two verification: the full migration chain on a real Postgres
Not a file-parsing claim this time. A throwaway
pgvector/pgvector:pg17container on its own port, seeded with.github/ci/test-db-bootstrap.sqland then every file insupabase/migrations/in order, exactly as.github/workflows/ci.ymldoes it. Container removed afterwards.all migrations applied, no error.Read back from that database:
Also confirmed against that database:
hive-embedding-defaultwith 3, which is the pre-existing exception the integration suite already records inpendingMultiRouteAliases.route-free-auto.supports_sttandsupports_ttsare still on the two healthy Groq audio routes. So/v1/batches,/v1/images/generations,/v1/images/editsand both voice endpoints still find an eligible route.policy_modemoved on no alias.hive-fastis stilllatency,hive-smallandhive-mediumstillpinned,hive-defaultstillstability,hive-autostillweighted. Only the route name insidefallback_orderchanged.internal/routing,internal/catalogandinternal/litellmconfigallokwith-tags integration.Running the whole
./apps/control-plane/...integration suite at once also produced failures inauditworker,marketplaceandtenants. Those are package-parallelism collisions on one shared database, not this change: each passes on its own against the same database, and none of them reads any of the four tables this branch touches.Two test-file notes a reviewer should see rather than discover
TestHiveFastIsPinnedToGroqAtCorrectedPriceis renamed toTestHiveFastIsPinnedToOneRouteAtItsUnchangedPriceand its provider and provider_model expectations updated, because the route legitimately moved. Its two price assertions (10500 and 42000) are unchanged and are now the DB-level guard that this repoint did not quietly reprice three aliases. Its comment records that these figures are no longer derivable from the upstream's cost, which is zero, and must not be "corrected" to match it.TestSelectRouteHiveFastResolvesToGroqAtGroqPriceinservice_test.gokeeps a name that no longer describes the catalog. It is a stub-driven test of theSelectRoutealgorithm with synthetic route fixtures; it reads neither the database nor any migration, so it is still a valid algorithm test. Left alone rather than renamed in an unrelated file. Push back if you would rather it were renamed.Residual question three, added by part two
There is now no chat alias on a second provider. With every free-model alias behind one 20-requests-per-minute endpoint and no fallbacks, a single upstream 429 takes the whole chat surface down for that minute. Options: (a) leave it, which is what the directive literally asks for; (b) keep one Groq chat route as a deliberate escape hatch, cheapest candidate being the deprecated
hive-fast, which costs almost nothing in Groq allowance and preserves a working alias to switch to; (c) fix #1089 first so at least the failure is a fast, correct 429 a client library can back off from. Recommendation: (c) then (b). Not done here because #1089's fix is not containable from this environment, per the reasoning above.