Skip to content

feat: route every chat alias to a free OpenRouter model, and halve hive-default and hive-auto - #1099

Merged
sakibsadmanshajib merged 2 commits into
mainfrom
feat/free-route-aliases-half-price
Aug 23, 2026
Merged

sakibsadmanshajib merged 2 commits into
mainfrom
feat/free-route-aliases-half-price

Conversation

@sakibsadmanshajib

@sakibsadmanshajib sakibsadmanshajib commented Aug 23, 2026 •

Copy link
Copy Markdown
Owner

Owner directive, 2026-08-23: route both hive-default and hive-auto to OpenRouter free models, and charge 50 percent of what is being charged now.

One migration plus one LiteLLM config edit. No Go change on the serving path.

First, a correction to the premise this task was handed with

The brief said hive-auto bills at actual upstream cost after PR #1012, and therefore had to be converted back to a fixed price with the "today" figure recovered from ledger rows.

That is not what #1012 did. It created a new alias, openrouter-auto, on a new route route-openrouter-auto-beta, in pricing_mode = 'upstream_actual'. Its own migration says so out loud: "litellm_model_name is deliberately NOT 'route-openrouter-auto': that name is the retired route id of the pre-existing hive-auto alias, which resolves to a completely different model."

hive-auto was never touched by it. Both aliases are plain fixed rows and have been since the 2026-08-22 restructure, so there is no pricing-mode conversion here and no need to reconstruct a price from ledger magnitudes. The old figures come from the statement that set them.

Before and after

Source for every old rate: supabase/migrations/20260822_02_catalog_alias_restructure.sql step 7.

alias unit old rate new rate source of old rate
hive-default credits per million prompt tokens 10500 5250 20260822_02 step 7
hive-default credits per million completion tokens 42000 21000 20260822_02 step 7
hive-default credits per million cache-read tokens (published, never billed) 0 0 20260822_02 step 7
hive-default credits per million cache-write tokens (published, never billed) 0 0 20260822_02 step 7
hive-auto credits per million prompt tokens 21000 10500 20260822_02 step 7
hive-auto credits per million completion tokens 84000 42000 20260822_02 step 7
hive-auto credits per million cache-read tokens (published, never billed) 0 0 20260822_02 step 7
hive-auto credits per million cache-write tokens (published, never billed) 0 0 20260822_02 step 7

Every figure halves with no remainder, so no rounding rule had to be invented and no float can enter. Both aliases stay pricing_mode = 'fixed' and price_unit = 'tokens'.

Why the cache columns are listed at 0 and not omitted: they are the only other price columns on the row, precedence.go never reads them, and both were already zeroed by 20260822_02 when the aliases moved to Groq routes declaring no cache support. Half of zero is zero, so they are asserted rather than changed. Naming them is the point: "every unit type the alias bills today" is answered exhaustively rather than by omission.

Why there is no Go change

apps/edge-api/internal/metering/precedence.go is the single implementation of the charge arithmetic (D-031). It builds exactly two UnitCharge values per request, prompt tokens at InputPriceCredits and completion tokens at OutputPriceCredits, sums them, divides once by a million and rounds half up. So halving those two columns halves every token charge exactly, for every request shape, with nothing else moving:

  • The credit hold does not move, and should not. For a fixed-price alias ReservationCredits returns the flat per-endpoint default (10000 for chat); it only raises a hold above that for an upstream_actual alias. The hold is an authorization, released in full at settlement.
  • The fail-closed path halves too. A missing usage block synthesises a token count from response bytes and prices it at the same catalog rate, so it lands at half the old figure rather than at zero or at the old rate.
  • Non-token modalities are untouched because they never read these columns. /v1/images/* reserves a hardcoded 5000 credits and internal/audio charges flat literals, both alias-independent (D-033, issue Audio billing ignores the catalog and charges flat hardcoded credits #627). See the residuals.
  • Which token classes are billed does not change. Only the rate does. There is a guard for this specifically, because a brief of this shape previously produced a 262x overcharge by widening what was billed.

The money proof: same request shapes, before and after

Produced by calling the production settlement function inference.CreditsForTokens directly, once with the old rates and once with the new ones, in the repo's own toolchain container. Not a reimplementation: that is the function both the streaming and the sync settlement paths call.

Two figures per shape. Exact is quantity times rate, summed, before the single division and round-half-up, which is where "exactly half" is provable. Settled is the whole-credit charge the ledger records, and both sides round independently.

alias request shape prompt completion old exact (credit-millionths) new exact exact ratio old settled new settled
hive-default typical chat turn 1200 400 29400000 14700000 0.500000 29 15
hive-default long context read 32000 800 369600000 184800000 0.500000 370 185
hive-default prompt only 1500 0 15750000 7875000 0.500000 16 8
hive-default completion only 0 900 37800000 18900000 0.500000 38 19
hive-default tool-call-only turn 850 40 10605000 5302500 0.500000 11 5
hive-default fail-closed byte estimate 72 1000 42756000 21378000 0.500000 43 21
hive-default one token each (floor) 1 1 52500 26250 0.500000 1 1
hive-default 24 completion tokens (floor boundary) 0 24 1008000 504000 0.500000 1 1
hive-default negative counts clamped -5 -5 0 0 n/a 0 0
hive-auto typical chat turn 1200 400 58800000 29400000 0.500000 59 29
hive-auto long context read 32000 800 739200000 369600000 0.500000 739 370
hive-auto prompt only 1500 0 31500000 15750000 0.500000 32 16
hive-auto completion only 0 900 75600000 37800000 0.500000 76 38
hive-auto tool-call-only turn 850 40 21210000 10605000 0.500000 21 11
hive-auto fail-closed byte estimate 72 1000 85512000 42756000 0.500000 86 43
hive-auto one token each (floor) 1 1 105000 52500 0.500000 1 1
hive-auto 24 completion tokens (floor boundary) 0 24 2016000 1008000 0.500000 2 1
hive-auto negative counts clamped -5 -5 0 0 n/a 0 0

Exactly half on every row, in integer rational arithmetic. Not "roughly half", not a fraction of a percent, not unchanged, and not zero.

Stated rather than smoothed over: the settled whole-credit charge can sit one credit above exactly half, because the old and new charges each round half up independently from exact rationals in a 2 to 1 ratio (29.4 rounds to 29 while 14.7 rounds to 15). The deviation is bounded at one credit, which is 0.00001 USD at the repo's 100000-credits-per-USD constant. The 1-credit floor in CreditsForTokens is the other bounded exception: it is what keeps a sub-credit request off zero, it predates this change, and it is why the two floor rows read 1 and 1. Both are visible in the table rather than hidden behind a "50 percent" claim the integers do not literally satisfy.

The replay enforced its own claims: it failed the run if any exact ratio was not exactly one half, if any new charge was zero where the old was positive, or if any new charge exceeded half by more than one credit.

Full capture, including the commands: docs/proof/free-route-aliases-half-price-2026-08-23/README.md.

The test that fails on the old rate

apps/control-plane/internal/routing/free_alias_pricing_test.go, offline, reusing the SQL parser already in that package. Ten guards, all positional: the migration declares its arithmetic in a -- HALVE| alias | field | old | new table, the test checks that table is arithmetically half, that its OLD column matches the rates actually in force (pinned here so a later edit to 20260822_02 cannot move the baseline), and then that the value each UPDATE assigns equals half. A correct comment above a wrong UPDATE is caught.

Mutation tested rather than asserted to work. Six mutations, six kills, control green:

mutation result
M1 revert hive-default input to the old 10500 FAIL TestFreeAliasPricesAreExactlyHalfTheOldRates
M2 hive-default output 21000 becomes hive-auto's 10500 FAIL TestFreeAliasPricesAreExactlyHalfTheOldRates
M3 route-free-auto loses supports_batch and both image flags FAIL TestFreeRouteAutoCarriesTheSoleCapabilityFlagsForward
M4 provider_model loses the :free suffix FAIL TestFreeAliasRoutesTargetAFreeOpenRouterModel
M5 route-free-default loses tools_supported FAIL TestFreeRoutesKeepTheCapabilitiesTheirAliasesServeToday
M6 the retired Groq routes left enabled FAIL TestRetiredGroqRoutesAreDisabledAndRepointed

M1 is the guard the brief asked for: red on the old rate, green on the new one.

The existing TestCatalogAliasPricesMatchProviderRates cannot cover these two prices, and that is structural rather than an oversight. It derives every credit figure as usd_per_million * 1.4 * 100000 (D-032), and its parseRate refuses a zero rate outright: "that is a mispricing, not a rate". A free upstream costs zero, so the margin formula yields zero, and zero is refused by SelectRoute as unpriceable. These two prices are owner-set, not cost-derived, and the halving relation is what replaces the formula as the checkable invariant.

Routing: which free model, chosen from live data

https://openrouter.ai/api/v1/models, fetched 2026-08-23: 422 models, 22 priced at zero on both prompt and completion. A literal openrouter/free id does exist; it is a router, not a model.

The capability bar comes from what these aliases serve today. Both current routes declare tools_supported = true, which is the column PR #206 routes tools, tool_choice and response_format on, plus supports_streaming and supports_reasoning. Of the 22 free models, five support all of tools, tool_choice, response_format and structured outputs. Joined to https://openrouter.ai/api/frontend/v1/all-providers:

free model provider trains on prompts retention
dots-studio/dots-3-note-preview:free AtlasCloud no retains, period not published
z-ai/glm-5.2:free Decart no zero retention
nvidia/nemotron-3-super-120b-a12b:free NVIDIA YES retains
nvidia/nemotron-nano-9b-v2:free NVIDIA YES retains
liquid/lfm-2.5-2.6b:free Liquid YES retains

A false green I shipped and then caught, recorded so the next person does not repeat it: the endpoints API reports NVIDIA's provider as Nvidia while the provider directory keys it under displayName NVIDIA. Joining on displayName alone misses the record and reports every NVIDIA free endpoint as no-training and zero-retention, the exact opposite of the truth. Join on both fields.

Live probes with the project's real key, and this is where the paper ranking fell apart:

  • z-ai/glm-5.2:free, the only zero-retention candidate and the strongest model on every published axis: 429 on four of four attempts, provider_error_code: upstream_429, limit_source: upstream_provider_shared_pool. That is Decart's shared pool, not our account's limit, and buying credits cannot raise it. Rejected on live evidence rather than on paper.
  • openrouter/free with no provider preference: five of five succeeded, but four landed on NVIDIA endpoints and two of five landed on nvidia/nemotron-3.5-content-safety:free, a moderation classifier, which answered a plain chat prompt with User Safety: safe and then with an empty string. Unfit for a chat alias on output quality alone, before the training question. This is also the exact shape issue hive-default and hive-auto route through OpenRouter's random free-model router, not the models their catalog prices are derived from #689 called a bug.
  • openrouter/free with provider: {data_collection: "deny"}: five of five, only Cohere and AtlasCloud, tool calls on five of five. The deny filter empirically excludes every NVIDIA endpoint and the classifier.
  • provider: {zdr: true}, the stricter form: 404, "No endpoints found matching your data policy (Zero data retention)". There is no zero-data-retention free endpoint reachable at all today.
  • dots-studio/dots-3-note-preview:free: 200 on a sync completion, on a tools request (a real tool_calls with correct arguments and finish_reason: tool_calls), on response_format: {"type":"json_object"} (valid JSON back) and on a streamed request. cost: 0 on every response.

Chosen: dots-studio/dots-3-note-preview:free for both aliases, pinned, with extra_body.provider.data_collection: deny and allow_fallbacks: false. It is the only free model that is simultaneously full parity, live verified working, and served by a provider that does not train on prompts. The deny preference is kept even though the model resolves to one provider today: if AtlasCloud's policy changes or a second provider appears, the request fails instead of quietly moving customer prompts to a provider that stores them.

Capability parity verdict

capability before (Groq gpt-oss) after (free model) verdict
tools, tool_choice declared, supported declared, live verified with a real tool call parity
response_format, structured outputs routed on tools_supported supported, live verified returning valid JSON parity
streaming declared declared, live verified, 20 chunks parity
reasoning / reasoning_effort declared declared, model enables reasoning by default parity
chat completions, completions, responses declared declared parity
embeddings false false unchanged
cache read, cache write false false unchanged, no cache rate published
batch, image generation, image edit on route-groq-auto only carried onto route-free-auto preserved, see below
vision (image input) none reachable upstream is text+image->text not declared, see below

Two notes on that table. provider_capabilities has no vision column, so the free model's image-input support is an undeclared property of the upstream rather than a new product claim; the 2026-08-22 note that no customer-reachable vision path exists is now inaccurate at the upstream level and the config comment says so. And the three media flags are a documented status-quo fiction inherited through two migrations: neither gpt-4.1-mini, nor gpt-oss-120b, nor this model generates images. They are carried because route-groq-auto is their sole carrier in the catalog, SelectRoute hard-filters on each flag, and batchstore sends NeedBatch = true for every batch, so disabling it without handing them on would leave /v1/batches, /v1/images/generations and /v1/images/edits with zero eligible routes for every alias in the system. M3 above is the guard.

Under-claiming is not a safe default here: both aliases are pinned to exactly one route, so matchesRequestedCapabilities drops the only candidate, SelectRoute returns ErrRouteNotEligible and writeRoutingError maps it to 422. On a pinned alias an under-claim is a failed request, not a withheld feature.

Rate limits, and what a user sees at the cap

Documented (https://openrouter.ai/docs/api_reference/limits.md, whose MDX constants resolve to real numbers): free variants are capped at 20 requests per minute, and at 50 per day below 10 dollars of lifetime credit purchases or 1000 per day at or above it. GET /api/v1/key reports is_free_tier: false for this account and project memory records a 10 dollar purchase, so 1000 per day is the expected tier. The key endpoint does not expose the purchased total, so that is an expectation, not a measurement; its own rate_limit field is documented as deprecated and returns requests: -1. Separately and more sharply, the Decart 429 shows a per-provider shared pool can refuse everything regardless of our standing.

What a customer sees at the cap today: a roughly 60 second wait and then a 502 whose message is context canceled. That is issue #1089, filed today and unchanged by this work. num_retries: 3 with request_timeout: 45 retries a rate-limited deployment three times, the SDK suites time out at 60 seconds, and the provider's own 429 with its retry hint never leaves the LiteLLM container log.

Not fixed here, and the reason is containment rather than appetite. The candidate fix (router_settings.retry_policy.RateLimitErrorRetries: 0) belongs in the litellm_settings block, which the config sync preserves verbatim on a live volume. A file edit there is inert on the box and cannot be verified from a developer machine with no SSH to it. #1089 reaches the same conclusion about itself and asks for a real reproduction. This change does make it materially more likely, since 20 requests per minute is much tighter than Groq's ceiling, so its priority moves from latent to likely.

Mitigation that does exist: hive-small and hive-medium stay on Groq and are now the only customer-reachable Groq chat routes besides the deprecated hive-fast. A rate-limited customer has a working alternative, but has to select it; nothing routes them there, because the gateway chat fallbacks were removed deliberately in 20260822_02 and re-adding one is the same inert-on-a-live-volume problem.

Logging and training terms

  • OpenRouter itself stores no prompts or completions unless the account opts in, and both opt-ins (private input/output logging, and letting OpenRouter use inputs/outputs for a 1 percent discount) are off by default. Request metadata (token counts, latency) is always retained.
  • AtlasCloud, serving the chosen model: training: false, retainsPrompts: true, no retention period published.
  • The route carries provider.data_collection: deny as a fail-closed guard.
  • The contradiction, written down because this product sells data sovereignty: a strict zero-retention posture is not achievable on any free OpenRouter endpoint today, proved by the zdr: true 404. The best available free posture is a provider that does not train but does retain for an unpublished period. The owner directed the move with this on the record.

Why new route ids rather than repointing the two Groq rows

The same reason 20260822_02 gave, and it applies in this direction too. The LiteLLM config sync merges field by field: the database owns only model, api_base and api_key, and every other key already on the entry survives so that hand-tuning sticks (mergeParams, issue #707). Retiring the route id makes the merge drop the whole stale entry, because a known route_id that is no longer active is deleted rather than updated. The rows are disabled, not deleted, so the change is reversible and SelectRoute filters them out.

price_class stays standard, identical to the routes being replaced. budget would arguably describe a free upstream better, but price_class feeds allow_price_class_widening, and keeping the same value means this repoint cannot change selection behaviour through a second mechanism.

The api_key follows automatically from providers.api_key_env for the row's provider slug, so openrouter resolves to OPENROUTER_API_KEY without the migration naming a secret. The doubled openrouter/ prefix on provider_model is correct: LiteLLM strips the leading one as its provider selector, exactly as openrouter/~deepseek/... and openrouter/openrouter/auto-beta already do. The trailing :free is load-bearing, not cosmetic: dropping it selects a paid endpoint of the same model, so the alias would charge a halved price against a real out-of-pocket cost and nothing else in the tree would notice. M4 is the guard.

Verification

  • go test ./apps/control-plane/... -count=1 -short: clean, no FAIL, no panic.
  • go test ./apps/edge-api/... -count=1 -short: clean, no FAIL, no panic.
  • New guards, verbose: 10 of 10 pass; 6 of 6 mutations kill a guard.
  • npm run lint:litellm-config: PASS. npm run lint:litellm-routing: PASS. npm run lint:proof-tokens: ok, 137 files scanned.
  • deploy/litellm/config.yaml parses and resolves 14 model entries, with both new routes carrying the intended extra_body.

Not proved, and it matters: that the running gateway serves the new route. The config is volume-seeded and the live change arrives through POST /internal/litellm/sync reading provider_routes. There is no SSH to the demo box from here and CI is the only remote hands, so the on-box confirmation belongs to the deploy run. deploy-demo-box.yml's "Assert model catalog prices agree with the model LiteLLM will call" step is the check that catches a stale volume, and its predicate (pricing_mode = 'upstream_actual' OR input_price_credits > 0) covers both of these rows. No ledger row at the new price exists yet either; that needs a served request after the migration applies.

No screenshot, because this change alters no UI surface. The console catalog table renders whatever price the API returns, and the two figures it will show are the ones proved above.

Residual questions for the owner

Both are stated rather than silently resolved, per the brief.

  1. hive-auto now costs exactly twice hive-default for the identical model. Both resolve to the same free upstream, so the surviving 2x gap buys the customer nothing. The directive fixes each alias at 50 percent of its own old price, which is what shipped; equalising them is a separate decision. Options: (a) leave as is, honest to the directive, indefensible to a customer who compares them; (b) equalise hive-auto to 5250 and 21000, a further reduction that cannot overcharge anyone; (c) give hive-auto a distinct larger free model. Recommendation: (b). Option (c) is blocked today: every other full-parity free model is served by a provider that trains on prompts, and the one zero-retention candidate returns 429 on every request.

  2. The image path is not halved. /v1/images/generations and /v1/images/edits reserve and settle a hardcoded 5000 credits that never reads the alias price, so an image request through hive-auto is billed the same as before. It is also not actually servable, since the upstream is a text model and those capability flags have been a documented fiction since 20260414_01. Options: (a) halve the literal to 2500, which halves it for every other alias too and so exceeds this directive's scope; (b) leave it and close the fiction under issue Audio billing ignores the catalog and charges flat hardcoded credits #627; (c) drop the flags, which deletes three endpoints catalog-wide. Recommendation: (b).

Buglog entry

{"id":"bug-2026-08-23-openrouter-provider-name-vs-displayname-join","date":"2026-08-23","title":"Joining OpenRouter endpoint provider_name against all-providers displayName silently reports a training provider as zero-retention","error_message":"A capability-and-data-policy scan of OpenRouter's free models reported every NVIDIA free endpoint as training:false, retainsPrompts:false, when NVIDIA's published policy is training:true, retainsPrompts:true","root_cause":"GET /api/v1/models/{slug}/endpoints reports the provider in its `name` form (`Nvidia`) while GET /api/frontend/v1/all-providers keys the record by `displayName` (`NVIDIA`). 18 of 82 providers have name != displayName. A dictionary keyed on displayName therefore misses the lookup entirely, and a default of an empty policy object reads as no-training and zero-retention rather than as unknown, so the failure presents as the safest possible answer instead of an error. The scan looked authoritative and was inverted on exactly the axis a data-sovereignty product cares about.","fix":"Key the provider lookup on both `name` and `displayName`, and treat a missing policy as UNKNOWN rather than as permissive. Re-ran the scan: only two of five full-capability free models are served by non-training providers, not four. Also verified the conclusion independently against live behaviour: with provider.data_collection deny set, five of five requests routed away from every NVIDIA endpoint, which agrees with the corrected join and contradicts the original one.","tags":["openrouter","data-policy","sovereignty","false-green","join-key-mismatch","provider-catalog"]}

Part two: the remaining Groq text routes, at prices deliberately unchanged

Second owner directive the same day, folded into this PR because it touches the same catalog rows and the same deploy/litellm/config.yaml, so a separate PR would conflict by construction.

Move the Groq TEXT and chat-completion models to OpenRouter free as well, to stop the Groq free-tier allowance being drained.

Shipped as a second migration, 20260823_21_groq_text_routes_to_openrouter_free.sql, deliberately separate from the repricing one. Splitting them is what makes "no price moves here" a structural property of a file rather than a claim in its header.

Prices for these three do NOT change, and that is the intended outcome

alias unit rate before rate after note
hive-small credits per million prompt / completion tokens 10500 / 42000 10500 / 42000 unchanged
hive-medium credits per million prompt / completion tokens 21000 / 84000 21000 / 84000 unchanged
hive-fast credits per million prompt / completion tokens 10500 / 42000 10500 / 42000 unchanged, deprecated alias

The owner's 50 percent instruction applies to hive-default and hive-auto only. Serving a same-priced alias from a free upstream widens margin, and that is intended. Stated explicitly so no later reader mistakes it for an oversight, and enforced two ways rather than promised:

  • TestGroqFreeRepointTouchesNoPrice fails if that migration assigns any of input_price_credits, output_price_credits, either cache column, pricing_mode or price_unit. The migration does not write model_aliases at all, so the guard holds in its strongest form.
  • The DB-level check reads the prices back after the whole chain applies, and they are byte for byte what they were.

hive-fast's cache columns stay at 1 and 4, the stale OpenRouter-era values 20260822_02 examined and deliberately left alone so a deprecated alias would not be repriced on any axis. Not touched here either, for the same reason.

Audio is out of scope, and the survey behind that

route-groq-stt (whisper-large-v3) and route-groq-tts (Orpheus) are untouched, and TestGroqFreeRepointLeavesAudioOnGroq fails if that migration so much as names them in an executable statement. GROQ_API_KEY stays required; what Groq no longer serves is chat.

"OpenRouter has no audio" would have been an overstatement, so here is the actual picture across all 422 models:

  • Text to speech: nothing usable. Not one model, free or paid, advertises supported_voices. The only entries with audio in their output modality are google/lyria-3-pro-preview and google/lyria-3-clip-preview (free, MUSIC generation, no voice selection) and openai/gpt-audio and openai/gpt-audio-mini (PAID, speech-to-speech chat). None is an OpenAI-compatible /v1/audio/speech endpoint, which is what internal/audio speaks. Orpheus has no replacement at any price.
  • Speech to text: three free models take audio as chat input (thinkingmachines/inkling:free, thinkingmachines/inkling-small:free, nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free). They are not a transcription endpoint: they are chat-completions models, while route-groq-stt is a LiteLLM mode: audio_transcription route that edge-api's audio handler forwards multipart audio to. Using one would be a new integration, not a repoint. All three are also served by providers whose published policy is training on prompts, a poor destination for dictated speech specifically.

Reported, not acted on, exactly as the directive asked. Bengali voice dictation (PR #1079) keeps working.

Capability parity, probed live because the parameter lists differ

This one genuinely needed checking rather than asserting. dots-3-note-preview lists tools, tool_choice, response_format, structured_outputs, reasoning, include_reasoning, max_tokens, temperature and top_p, and does not list reasoning_effort, stop, frequency_penalty, presence_penalty, seed, top_k or logprobs. The Groq gpt-oss models it replaces accept several of those. So: does an unlisted parameter fail, or is it ignored?

Twelve request shapes, live:

request shape result
baseline 200
reasoning_effort=low 200
reasoning_effort=high 200
stop 200
frequency_penalty 200
presence_penalty 200
seed 200
top_k 200
logprobs + top_logprobs 200
n=2 200
response_format json_schema, strict 200
reasoning_effort + provider.require_parameters 200

All twelve returned 200 with a correct answer. There is no request shape that works today and fails after the repoint, which is the regression this check exists to rule out. The require_parameters case is the interesting one: OpenRouter considers reasoning_effort satisfied by this endpoint even under strict parameter enforcement.

Two behavioural differences that are not failures, recorded so nobody files them as new bugs:

  • reasoning_effort is accepted and never rejected, but does not reliably modulate effort: low produced more reasoning tokens than high in the same run. Requests keep working; the knob stops being meaningful.
  • n=2 returns one choice. That is OpenRouter's existing behaviour for a provider that does not implement n, identical on the routes being replaced.

Capability FLAGS are carried across per route rather than uniformly. route-free-small and route-free-medium mirror their Groq originals including supports_reasoning = true; route-free-fast keeps supports_reasoning = false, which is status-quo preservation of an under-claim route-groq-fast has carried since its original 20260331_02 seed and that 20260822_02 examined and deliberately left alone. Widening it would be safe, but a routing migration is not where a deprecated alias should quietly gain a feature.

None of these three is a sole carrier of a media flag, so disabling them removes no endpoint. TestRepointedGroqTextRoutesKeepTheirCapabilities asserts both directions: the flags they must have, and that they claim none of the three media flags that belong on route-free-auto.

Concentration risk, stated rather than buried

After this change every customer-reachable chat alias except the two paid DeepSeek ones resolves to ONE model at ONE provider on ONE free endpoint. Five aliases: hive-default, hive-auto, hive-small, hive-medium, hive-fast. There are no gateway fallbacks, by deliberate design (20260822_02 removed them because a fallback answers from a model the alias was not priced against).

So the failure mode moves rather than disappearing. OpenRouter documents free variants at 20 requests per minute, and 50 or 1000 per day depending on lifetime credit purchases; that per-minute ceiling is tighter than the Groq daily allowance this was meant to escape. Separately, the Decart 429 recorded in part one shows an upstream provider shared pool can refuse everything regardless of our account standing, and buying credits cannot raise it.

It also removes the mitigation part one could still point at. An hour ago a rate-limited customer could select hive-small and land on a different provider. There is no such alias now.

What a demo user sees at the cap is unchanged: a roughly 60 second wait and then a 502 whose message is context canceled, never the provider's own 429 with its retry hint. That is issue #1089. Not fixed here for the containment reason given in part one: the candidate fix lives in the litellm_settings block the config sync preserves verbatim, so a file edit is inert on a live box and cannot be verified from a developer machine with no SSH to it. This change raises that issue's priority from latent to likely.

This does not close issue #1088. That is CI consuming the live demo's provider allowance, a different cause with a different owner, and there is deliberately no Fixes #1088 line anywhere in this PR.

Part two verification: the full migration chain on a real Postgres

Not a file-parsing claim this time. A throwaway pgvector/pgvector:pg17 container on its own port, seeded with .github/ci/test-db-bootstrap.sql and then every file in supabase/migrations/ in order, exactly as .github/workflows/ci.yml does it. Container removed afterwards. all migrations applied, no error.

Read back from that database:

   alias_id   | in_credits | out_credits |     route_id       |  provider  |                provider_model
--------------+------------+-------------+--------------------+------------+-------------------------------------------------
 hive-auto    |      10500 |       42000 | route-free-auto    | openrouter | openrouter/dots-studio/dots-3-note-preview:free
 hive-default |       5250 |       21000 | route-free-default | openrouter | openrouter/dots-studio/dots-3-note-preview:free
 hive-fast    |      10500 |       42000 | route-free-fast    | openrouter | openrouter/dots-studio/dots-3-note-preview:free
 hive-medium  |      21000 |       84000 | route-free-medium  | openrouter | openrouter/dots-studio/dots-3-note-preview:free
 hive-small   |      10500 |       42000 | route-free-small   | openrouter | openrouter/dots-studio/dots-3-note-preview:free
 hive-stt     |          0 |     4316667 | route-groq-stt     | groq       | groq/whisper-large-v3
 hive-tts     |          0 |     3080000 | route-groq-tts     | groq       | groq/canopylabs/orpheus-v1-english

Also confirmed against that database:

  • One enabled route per alias. The violations query returns exactly one row, hive-embedding-default with 3, which is the pre-existing exception the integration suite already records in pendingMultiRouteAliases.
  • The three batch and image flags have a healthy carrier, route-free-auto. supports_stt and supports_tts are still on the two healthy Groq audio routes. So /v1/batches, /v1/images/generations, /v1/images/edits and both voice endpoints still find an eligible route.
  • policy_mode moved on no alias. hive-fast is still latency, hive-small and hive-medium still pinned, hive-default still stability, hive-auto still weighted. Only the route name inside fallback_order changed.
  • Idempotent. Both new migrations re-applied to the same database a second time, both clean, prices identical afterwards.
  • Integration suite green against that database: internal/routing, internal/catalog and internal/litellmconfig all ok with -tags integration.

Running the whole ./apps/control-plane/... integration suite at once also produced failures in auditworker, marketplace and tenants. Those are package-parallelism collisions on one shared database, not this change: each passes on its own against the same database, and none of them reads any of the four tables this branch touches.

Two test-file notes a reviewer should see rather than discover

  • TestHiveFastIsPinnedToGroqAtCorrectedPrice is renamed to TestHiveFastIsPinnedToOneRouteAtItsUnchangedPrice and its provider and provider_model expectations updated, because the route legitimately moved. Its two price assertions (10500 and 42000) are unchanged and are now the DB-level guard that this repoint did not quietly reprice three aliases. Its comment records that these figures are no longer derivable from the upstream's cost, which is zero, and must not be "corrected" to match it.
  • TestSelectRouteHiveFastResolvesToGroqAtGroqPrice in service_test.go keeps a name that no longer describes the catalog. It is a stub-driven test of the SelectRoute algorithm with synthetic route fixtures; it reads neither the database nor any migration, so it is still a valid algorithm test. Left alone rather than renamed in an unrelated file. Push back if you would rather it were renamed.

Residual question three, added by part two

There is now no chat alias on a second provider. With every free-model alias behind one 20-requests-per-minute endpoint and no fallbacks, a single upstream 429 takes the whole chat surface down for that minute. Options: (a) leave it, which is what the directive literally asks for; (b) keep one Groq chat route as a deliberate escape hatch, cheapest candidate being the deprecated hive-fast, which costs almost nothing in Groq allowance and preserves a working alias to switch to; (c) fix #1089 first so at least the failure is a fast, correct 429 a client library can back off from. Recommendation: (c) then (b). Not done here because #1089's fix is not containable from this environment, per the reasoning above.

…half price

Owner directive, 2026-08-23: route both aliases to OpenRouter free models and
charge 50 percent of the current price.

Prices halve exactly, with no remainder and no rounding decision: hive-default
goes from 10500 in and 42000 out to 5250 and 21000, hive-auto from 21000 and
84000 to 10500 and 42000, both in credits per million tokens. The old figures
come from 20260822_02 step 7, the statement that set them.

Both aliases route to dots-studio/dots-3-note-preview:free, which is the only
zero-priced model on OpenRouter today that is simultaneously full parity on
tools, tool_choice, response_format and structured outputs, live verified
working, and served by a provider that does not train on prompts. The route
carries provider.data_collection deny as a fail closed guard.

No Go change on the serving path. precedence.go bills prompt and completion
tokens from the two alias price columns in one place, so halving those columns
halves every charge exactly, including the byte estimated fail closed shape, and
never to zero. Which token classes are billed does not move.

route-free-auto carries supports_batch, supports_image_generation and
supports_image_edit forward from route-groq-auto, which is their sole carrier in
the catalog. Without that, disabling it would leave /v1/batches,
/v1/images/generations and /v1/images/edits with zero eligible routes for every
alias in the system.
…ces unchanged

Second half of the same owner directive: move the Groq TEXT and chat completion
models to OpenRouter free to stop the Groq free tier allowance being consumed.

hive-small, hive-medium and hive-fast are repointed to
dots-studio/dots-3-note-preview:free. Their customer facing prices do NOT move.
The 50 percent instruction applies to hive-default and hive-auto only, so
serving these three from a free upstream simply widens margin, which is the
intended outcome. That is enforced rather than promised:
TestGroqFreeRepointTouchesNoPrice fails if this migration writes any price
column, and the migration writes to model_aliases not at all.

Groq speech to text and text to speech are untouched. A survey of all 422
OpenRouter models found no model at all advertising selectable voices and no
OpenAI compatible speech endpoint, so Orpheus has no replacement at any price,
and the only free models accepting audio take it as chat input rather than
through a transcription endpoint. GROQ_API_KEY stays required.

Capability parity was probed live rather than inferred, because the free model
does not list reasoning_effort, stop, seed, the penalties, top_k or logprobs.
Twelve request shapes covering all of those plus json_schema structured output
returned 200, so no request shape that works today fails after the repoint.

Concentration risk, accepted and recorded: every chat alias except the two paid
DeepSeek ones now resolves to one model at one provider on one free endpoint
capped at 20 requests per minute, with no gateway fallbacks. The failure mode
moves rather than disappearing, and the workaround of switching to hive-small on
a different provider no longer exists. What a customer sees at the cap is issue
#1089, unchanged here. This does not close issue #1088.
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@coderabbitai

coderabbitai Bot commented Aug 23, 2026

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 45 minutes.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: e56f410b-c451-4254-8451-b3aa118d353d

📥 Commits

Reviewing files that changed from the base of the PR and between efd507a and 5227171.

📒 Files selected for processing (7)
  • .env.example
  • apps/control-plane/internal/routing/catalog_pricing_integration_test.go
  • apps/control-plane/internal/routing/free_alias_pricing_test.go
  • deploy/litellm/config.yaml
  • docs/proof/free-route-aliases-half-price-2026-08-23/README.md
  • supabase/migrations/20260823_20_free_route_aliases_half_price.sql
  • supabase/migrations/20260823_21_groq_text_routes_to_openrouter_free.sql

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sakibsadmanshajib
sakibsadmanshajib merged commit 74602b3 into main Aug 23, 2026
25 checks passed
@sakibsadmanshajib
sakibsadmanshajib deleted the feat/free-route-aliases-half-price branch August 23, 2026 20:11
sakibsadmanshajib added a commit that referenced this pull request Aug 23, 2026
…ct (#1061)

## What this is

A machine gate that runs after `deploy-demo-box` succeeds and asserts
that the product actually works. A deploy exiting 0 is not that claim:
PR #787 merged green and took chat sign-in down, and three consecutive
deploys failed at `migrate` while the box silently stayed on
two-week-old code.

No human approval anywhere. There is no `environment:` key on any job,
and there must never be one: `main` went `7cfc460ea` (added an approval
gate) then `09607ff05` (removed it). The rigour comes from stronger
checks, not from a person in the loop.

## What it gates on

| # | Check | What would otherwise pass |
| --- | --- | --- |
| 1 | **Chat answers.** `GET /api/config` returns 200 with `status:
true` and an `oauth.providers.oidc` entry. | The static shell serves a
full 200 HTML page over a dead backend. A page title or a bare status
code passes straight through an outage. |
| 2 | **Sign-in completes.** A minted session must resolve at `GET
/api/v1/viewer` to the **same user id** it was minted for. | A rendered
login page proves nothing. Control-plane revalidates every bearer
against GoTrue per request, so this cannot be faked; comparing the id
rather than the status rejects a 200 carrying the wrong identity. |
| 3 | **A ledger entry lands.** A real completion through a freshly
minted key must produce a new `usage_charge` row with `credits_delta <
0`. | Billing failing open is invisible from every other angle: the
request succeeds and the current trace would be a row that is not there.
That is exactly what happened for three days in July 2026. |

Plus a **preflight**, which runs first and gates too: does the Supabase
configuration in front of us name a project that is alive, with keys
that belong to it.

And one **measurement that reports without gating**: whether every
container in the stack is running the image its tag currently resolves
to (issue #869). Why it reports rather than gates is in the next
section.

The key for check 3 is minted **through the session**, not taken from a
pre-provisioned key, so key and ledger are guaranteed to share an
account. It is revoked in a `finally` block, and a revocation that
itself fails is now a check failure rather than a warning.

## Honest coverage: the three failure modes this was asked to catch

**A deploy that reports a rebuild while the container keeps serving the
old binary (#869): measured and reported, NOT gated.** The comparison is
sound where it matches: 15 of 20 containers match their tag exactly, and
control-plane, edge-api and open-webui moved from mismatched to matched
across a deploy that recreated them, so the check tracks real state.
What the image-id comparison cannot yet establish on this box is the
**direction**, which is the entire difference between a stale binary and
harmless drift: five containers report an image id that `docker image
inspect` cannot resolve at all while their tag resolves to an image
built seconds before the container was created, and a container newer
than its tag's image did not miss that image. Failing a deploy on that
would leave this gate red on every deploy from the day it merges, and a
gate that is always red is the one nobody reads on the day it means
something. So the numbers ship with the run, the evidence goes to #869,
and turning it fatal is a one-line change with a comment saying exactly
that. **This gate does not currently catch #869.**

**A deploy that is green while shipping a dead URL (#1059 class):
caught, and proved.** This is the preflight. It was, however, weaker
than its own header claimed: pointed at `https://chat-hive.scubed.co`, a
single-page app that answers 200 with HTML for every path it does not
know, it reported the Supabase configuration sound. Measured on the live
box, not imagined. It now requires the `/auth/v1/health` body to be
GoTrue's own health document, and the `config` negative control
reproduces the failure exactly that way, since a hostname updated in one
place and left behind in another is the shape #1059 actually took.

**A migration that fails while the tracker reports success: NOT caught
here, and it should not be.** That check already exists upstream and
closer to the event: `scripts/apply-migrations.sh` plus
`scripts/probe-applied-migrations.py` decide the baseline state from the
schema itself rather than from the ledger, precisely because an empty
ledger has two opposite meanings, and `migrate` runs before `deploy` so
a failure stops the deploy. This gate notices such a failure only if it
changes chat, sign-in or billing behaviour. The gap worth closing
separately is a post-deploy drift assertion (re-run the probe read-only
and fail on any applied-versus-present disagreement); it is deliberately
not in this pull request.

## Where each input comes from

URLs and keys come from the **box**. Six workflows read `SUPABASE_*`
from GitHub secrets naming a hosted project that has been deleted, and
nothing went red (#1059). The box's `.env` is what the running
containers actually use.

Two corrections to that were needed before this could run at all:

- `SUPABASE_URL` on the box is `http://caddy-supabase`, an in-network
compose hostname. This job runs on the host, outside the compose
network, so signing in against it dies in DNS. The auth origin is taken
from `NEXT_PUBLIC_SUPABASE_URL` (the public origin the product's own
browser clients use), with the anon key read from
`NEXT_PUBLIC_SUPABASE_ANON_KEY` first so the origin and the key stay
paired. A single-label host now fails immediately, naming why.
- Cloudflare fronts both demo hostnames and answers 403 "Error 1010" to
the literal `Python-urllib/3.x` User-Agent. Measured: the same `GET
/api/config` returns 403 with that default and 200 with a named agent,
from the same host in the same second. Both scripts send one.

All three verification targets (`HIVE_CHAT_URL`,
`HIVE_CONTROL_PLANE_URL`, `HIVE_EDGE_API_URL`) are required by the
script from its environment with no default: a checker that silently
assumes a fixed demo host can return green for a product nobody
deployed. The workflow exports them explicitly, box `.env` first, with
the deployment's public hostnames as a reviewed fallback whose source
every run reports.

The verification **identity** comes from the box's `.env`. That is not a
contradiction. A stale URL is dangerous because nothing exercises it; a
stale credential cannot hide, because signing in with it is the next
thing this job does.

## Verification identity: configured, provisioned, proved

Resolved during this session; nothing is left for the owner here.

- `HIVE_VERIFY_EMAIL` and `HIVE_VERIFY_PASSWORD` exist in
`/home/sakib/hive/.env` on the box (added outside this branch; presence
verified by reading the file's key names only).
- The account behind them was checked read-only against the database: a
real GoTrue user, email confirmed, single tenant membership, own billing
account.
- It was not usable as first written: membership was MEMBER and balance
zero, which the gate itself surfaced as a 403-shaped failure class and
an upstream `insufficient_quota`. Provisioning applied two scoped writes
limited to that user's own rows: membership raised to OWNER (minting an
API key needs `api_keys:write`, granted to OWNER only), and one
idempotent grant entry of 10,000 credits on its own billing account
(`idempotency_key post-deploy-verify-seed-2026-08-23`, unique index
makes reruns inert).
- First fully-green verification followed immediately, on the box, all
three checks passing including a real charge of `-1` credit and clean
key revocation.

## Negative controls

Every gating check has one, and each removes the guarded condition for
real:

| Value | What it removes | Expected |
| --- | --- | --- |
| `chat` | reads the static shell at `/` instead of the backend config |
red |
| `signin` | flips every bit of the bearer's first decoded signature
byte, header and payload intact | red |
| `ledger` | skips the completion, so nothing bills | red |
| `config` | points the auth origin at another live host in this
deployment | red |
| `freshness` | compares against an image id nothing can be running |
red |

No value can make anything pass; it can only cause a failure. Selectable
from `workflow_dispatch` after this merges, and from a
`verify-negative:<value>` label before then. Selecting `signin` or
`ledger` with no identity configured fails immediately naming that
precondition rather than surfacing later as a confusing argparse error.

## Verification, run by run

All against the live box unless noted.

| Run / evidence | What it shows |
| --- | --- |
| On-box full run, 2026-08-23 ~20:42 UTC | **Full green baseline, all
three checks**: chat backend config 200 with `status: true` and OIDC
provider; minted session resolved at the control-plane to the same user;
real completion answered 200, produced `usage_charge`
`credits_delta=-1`, key revoked. |
|
[32665274872](https://github.com/sakibsadmanshajib/hive/actions/runs/32665274872)
| `signin` control red through the assertion: corrupted decoded
signature bytes, control-plane answered 401 `invalid or expired token`.
|
|
[32665392330](https://github.com/sakibsadmanshajib/hive/actions/runs/32665392330)
| `ledger` control red through the assertion: completion skipped, no new
`usage_charge` within 90s, reported as billing failing open. |
|
[32660681673](https://github.com/sakibsadmanshajib/hive/actions/runs/32660681673)
| Earlier partial green baseline: preflight passed from GoTrue v2.189.0
at the public origin, chat check passed, 20 containers compared, missing
identity reported as unconfigured coverage. |
|
[32659829693](https://github.com/sakibsadmanshajib/hive/actions/runs/32659829693)
| `chat` control red, and for the right reason: `/` answered 200 with
HTML, so the check names the static shell rather than the backend. |
|
[32660319162](https://github.com/sakibsadmanshajib/hive/actions/runs/32660319162)
| `config` control red: `GET /auth/v1/health` answered 200 from a host
that is not GoTrue, and the preflight says so instead of passing. |
|
[32660549692](https://github.com/sakibsadmanshajib/hive/actions/runs/32660549692)
| `freshness` control red through the assertion path, every container
reported and the step failing on the comparison. |
|
[32660408662](https://github.com/sakibsadmanshajib/hive/actions/runs/32660408662)
| The same control failing for the wrong reason before it was fixed.
Kept because a control that fails for the wrong reason proves nothing,
and this is what that looks like. |
|
[32634926680](https://github.com/sakibsadmanshajib/hive/actions/runs/32634926680)
| The original failure this PR had to fix: no verification identity in
the box `.env`. |
|
[32664630991](https://github.com/sakibsadmanshajib/hive/actions/runs/32664630991)
| The new response-shape guard catching a real condition live: a genuine
zero-row account answers `{"entries": null}`, which is the Go handler's
designed empty answer; the scan now reads null as empty and any other
non-list type still fails naming its type. |
|
[32665543894](https://github.com/sakibsadmanshajib/hive/actions/runs/32665543894),
[32665768538](https://github.com/sakibsadmanshajib/hive/actions/runs/32665768538)
| Upstream capacity noise, documented not hidden: after #1099 every chat
alias routes to an OpenRouter free model, whose per-key daily capacity
intermittently answers 429 `insufficient_quota` while chat itself stays
up. The completion is retried up to `HIVE_VERIFY_COMPLETION_ATTEMPTS`
(default 3) with `HIVE_VERIFY_RETRY_DELAY` (default 20s) between
attempts; exhausted retries fail naming the last answer. Expect the
ledger check to flap with OpenRouter free-tier capacity until the route
mix changes. |

## Review

- **CodeRabbit CLI**: ran twice (round 1: three findings, three fixes;
round 2 on the new commits: four findings, four fixes: target URLs
required rather than defaulted, JSON shapes validated before field
access, byte-changing signature mutation, revocation failure fatal,
secret lengths suppressed for password/service-role values,
deterministic negative-control invocation).
- **`ecc:code-review`**: ran as a structured seven-category pass over
the three changed files at head revision. Findings folded back into the
fixes above plus one consistency fix (service-role key length suppressed
in the preflight output). Posted as a COMMENT review; the builder does
not approve its own PR.
- **Plain adversarial pass**: ran, against the live box rather than by
reading. Produced the Cloudflare user-agent block, the non-GoTrue
health-document gap, subset-run summary wording, notifier scope,
freshness-control ordering, anon-key pairing, the entries-null
discovery, and the 429 retry semantics.
- **Mandatory security review on an auth and money path**: manual pass
covered: no value printed anywhere (presence only for password and
service-role values, masked on export), no credential crossing
`$GITHUB_ENV`, a step summary or a job boundary, no password ever
written or rotated (`docs/live-test-auth.md`), the minted key revoked in
a `finally` with failure now fatal, fork head repositories excluded from
the self-hosted runner, `persist-credentials: false`, and the real spend
bounded (one completion, `max_tokens: 64`, dedicated identity, never the
demo account).
- **`/codex:adversarial-review`**: SKIPPED, capability gap, not a clean
pass. Codex usage limits were reached earlier today (recorded on the PR
timeline); the stream could not run.

## Failure handling: alert, not rollback

Deliberately no automatic rollback. `migrate` runs before `deploy` and
forward migrations are not reversible by redeploying code, so an
automatic rollback would leave old code against a new schema, which is
worse than a bad deploy.

The notifier uses its own `ci-failure:post-deploy-verify` label rather
than sharing `ci-failure:deploy-demo-box`, since this workflow cannot
join that job's `needs` list and its ledger check can go red for reasons
unrelated to a deploy. It fires **only on `workflow_run`**: issue #1062
exists because a pull request run filed it, and every negative-control
run is expected to fail. It dedupes by exact title over the issues list
API rather than the eventually consistent search API, and
`continue-on-error` plus `|| true` mean it can never alter the verdict.

## Not a user-visible surface

No screenshot: this changes a workflow and two scripts. Nothing in the
product's UI moves, so the visual-proof rule does not apply. The
evidence is the run table above.

## Test plan

- [x] Green baseline: the checks that can run pass against the running
box (partial CI baseline, then a full three-check green run on the box
once the identity was provisioned).
- [x] `negative_control=chat` goes red.
- [x] `negative_control=config` goes red, restore goes green.
- [x] `negative_control=freshness` goes red through the assertion.
- [x] `negative_control=signin` goes red through the assertion (run
32665274872).
- [x] `negative_control=ledger` goes red through the assertion (run
32665392330).
- [x] Preflight fails legibly on a Supabase configuration that answers
HTTP but is not our project.
- [x] The notifier cannot fire from a pull request run.
- [x] CodeQL alert 27 cleared, thread resolved.
- [ ] One more fully-green CI run once OpenRouter free-route capacity
allows a completion through; the on-box full run already proves the path
end to end.

## Buglog entry

To be appended to `.wolf/buglog.jsonl` on `main` in a buglog-only pull
request after this merges, per issue #873.

```json
{"date":"2026-08-23","title":"Cloudflare answered 403 Error 1010 to the default Python urllib user agent, so every HTTP check against the demo hostnames failed","error_message":"HTTP 403 {\"title\":\"Error 1010: Access denied\",\"detail\":\"The site owner has blocked access based on your browser's signature\"}","root_cause":"urllib sends User-Agent: Python-urllib/3.x by default, and the Cloudflare bot rules in front of chat-hive.scubed.co and console-hive.scubed.co block that signature. curl from the same host in the same second returned 200, which is why the failure read as a dead backend rather than a blocked client.","fix":"scripts/post-deploy-verify.py and scripts/preflight-supabase-config.py set an explicit named User-Agent on every request.","tags":["cloudflare","http","verification","false-negative","demo-box"]}
{"date":"2026-08-23","title":"The Supabase preflight reported a wrong host as a healthy project because it only checked the status code","error_message":"PREFLIGHT PASSED against SUPABASE_URL=https://chat-hive.scubed.co","root_cause":"The liveness check asserted GET /auth/v1/health returned 200 and nothing about the body, while chat-hive.scubed.co is a single-page app that answers 200 with HTML for every unknown path.","fix":"The preflight requires the health body to be GoTrue's own document and fails naming what answered instead.","tags":["supabase","preflight","false-positive","issue-1059"]}
{"date":"2026-08-23","title":"post-deploy-verify filed an issue claiming the demo box was broken from a pull request run","error_message":"issue #1062 post-deploy-verify is failing against the demo box, verify=failure, trigger=pull_request","root_cause":"The report-failure job was gated on failure() alone, so a label-gated pull request run of the workflow, including every negative-control run that is expected to fail, filed an outage issue for the deployment.","fix":"The notifier requires github.event_name == 'workflow_run', and a missing verification identity files under its own title.","tags":["github-actions","notifier","false-alarm","issue-1062"]}
{"date":"2026-08-23","title":"The verification could not sign in at all because the box's SUPABASE_URL is an in-network compose hostname","error_message":"SUPABASE_URL=http://caddy-supabase does not resolve from the runner host","root_cause":"Services inside the compose network reach auth at http://caddy-supabase; this job runs on the host outside that network. The public auth origin is NEXT_PUBLIC_SUPABASE_URL, a different value on this box.","fix":"The workflow reads NEXT_PUBLIC_SUPABASE_URL first, pairs it with NEXT_PUBLIC_SUPABASE_ANON_KEY, and fails immediately when the resolved host has no dot in it.","tags":["supabase","demo-box","networking","verification"]}
{"date":"2026-08-23","title":"The verifier crashed with AttributeError instead of a named failure when a 200 carried the wrong JSON shape","error_message":"AttributeError on payload.get / viewer.get / entries iteration","root_cause":"Decoded JSON was field-accessed without type validation, and main() catches CheckFailed only, so a malformed 200 skipped the per-check summary entirely.","fix":"Sign-in, viewer and ledger payloads are shape-checked before access and raise the named CheckFailed.","tags":["python","verification","error-handling"]}
{"date":"2026-08-23","title":"A real zero-row ledger answered {\"entries\": null} and the new shape guard rejected a correct API answer","error_message":"ledger read answered 200 but carried entries as NoneType, expected a list","root_cause":"The control-plane handler builds the list with var entries []LedgerEntry, so an empty account marshals as null, not []; the guard written for malformed bodies treated the designed empty answer as malformed. Found live on the first full run against the box.","fix":"Null entries scan as the empty page; any other non-list type still fails naming its type.","tags":["go","json","verification","null-semantics"]}
{"date":"2026-08-23","title":"Ledger check flapped on upstream 429 after the chat aliases moved to an OpenRouter free route","error_message":"the completion answered 429 insufficient_quota on attempts 1..3 while an identical completion minutes earlier returned 200 and charged normally","root_cause":"#1099 repointed every chat alias at an OpenRouter free model whose per-key daily capacity intermittently refuses service; the verifier's claim is that a served request bills, not that an upstream always serves.","fix":"Bounded retries (HIVE_VERIFY_COMPLETION_ATTEMPTS=3, HIVE_VERIFY_RETRY_DELAY=20s) treat 429/5xx as retryable; exhausted retries fail naming the last answer. Documented as expected flap until the route mix changes.","tags":["openrouter","rate-limit","verification","flaky-gate"]}
```


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
- Added automated post-deployment verification after successful
deployments, with support for manual runs and eligible pull requests.
- Verifies chat availability, sign-in, billing completion, container
freshness, and deployment configuration.
- Supports targeted checks, negative testing, retries, polling, and safe
credential handling.
- Added Supabase configuration preflight checks for environment
settings, tokens, URLs, and service health.
- Reports skipped checks when credentials are unavailable without
incorrectly failing verification.

- **Bug Fixes**
- Added deduplicated issue notifications for verification failures or
missing coverage without changing verification results.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

## Billing test evidence (CodeRabbit follow-up)

- `go test ./apps/edge-api/... -count=1 -short` in the toolchain
container: all packages pass, including the reservation guard tests
covering the 429 mapping (`reservation_guard`) and the strict-mode 409
to 429 translation.
- `go test ./apps/control-plane/... -count=1 -short`: all packages pass,
including the reservations policy tests for the flat hold and
PolicyError path (`service.go` strict mode).
- Live evidence: run 32669144332 rerun completed success at
2026-08-23T23:10:22Z with the full three-check green (chat, signin,
ledger), after the verify workspace balance was topped up past the flat
10000 credit hold. The named preflight added in stacked PR #1105 makes
this failure mode self-diagnosing going forward.
sakibsadmanshajib added a commit that referenced this pull request Aug 23, 2026
…1105)

Stacked on #1061 (targets its branch, since
scripts/post-deploy-verify.py exists only there). Not for merge until
#1061 lands, then rebase or retarget to main.

## What happened live (2026-08-23)

The post-deploy-verify ledger check failed every run with:

```
HTTP 429 {"error":{"message":"You exceeded your current quota, please check your plan and billing details.","type":"insufficient_quota","code":"insufficient_quota"}}
```

The wording reads like an upstream provider quota error. It is not.
Every upstream layer tested clean: direct OpenRouter 200, LiteLLM
route-deepseek-v4-flash 200, catalog alias mapping correct, zero
requests reaching litellm during the failures.

## Root cause

1. The verify workspace was seeded with exactly 10000 credits
(`post-deploy-verify-seed-2026-08-23` grant).
2. The first passing run consumed 1 credit, leaving 9999.
3. Every chat completion reserves a flat **10000 credit hold** before
dispatch (`apps/edge-api/internal/inference/chat_completions.go` passes
that endpoint default to `ReservationCredits`, which falls back to it
because `reservation_estimate_credits` is null on every fixed-price
alias).
4. Control-plane refuses the reservation with a strict-mode PolicyError
409 (`reservation exceeds available credits`,
`apps/control-plane/internal/accounting/service.go`).
5. Edge maps that 409 to HTTP 429 `insufficient_quota` with the
OpenAI-style message text
(`apps/edge-api/internal/inference/reservation_guard.go`). No
usage_events row is written because nothing was ever dispatched.

Refused requests consume nothing, so the balance never recovers and
every rerun fails identically. The timing correlation with #1099's
deploy was coincidence.

## What this PR changes

A named preflight in `check_ledger` before any spend: read `GET
/api/v1/accounts/current/credits/balance` and fail with the real reason
plus remediation when available credits are below the flat hold. A
non-200 read warns and proceeds so the completion attempt still speaks
for itself; the negative control skips both spend and preflight.

## Evidence

- Reproduced through the public edge with a key on the verify workspace
at balance 9999: exact 429 body above.
- Fixed operationally by a +10000 grant ledger row
(`post-deploy-verify-topup-2026-08-23`); same call then answered 200.
- Rerun of post-deploy-verify run 32669144332 completed success after
the top-up.

## Buglog entry

{"error_message":"post-deploy-verify ledger check fails with HTTP 429
insufficient_quota from /v1/chat/completions","root_cause":"verify
workspace balance 9999 credits fell below the flat 10000 credit chat
reservation hold after the first passing run consumed one credit;
control-plane answers 409 policy rejection and edge maps it to a 429
whose message mimics an upstream provider quota
error","fix":"operational +10000 credit grant to the verify workspace
ledger; PR adds a named balance preflight so the gate reports this
condition instead of an upstream-looking
429","tags":"billing,reservation,fail-closed,post-deploy-verify,misleading-error"}

## Test plan

- [x] py_compile passes
- [x] Live reproduction of the 429 at balance 9999 and 200 after top-up
- [x] Gate rerun green against the topped-up box
- [ ] Negative path: balance below hold produces the new named
CheckFailed (exercisable by draining the workspace again)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

A provider rate limit reaches the customer as a 60 second wait and then 502 context canceled, not as a 429

1 participant