fix: wire read-aloud TTS to the gateway instead of a dead default - #1079
Conversation
Open WebUI's text-to-speech was never wired to the gateway (#997): every audio.tts.* key is persistent config, so a first boot seeded engine="" (bundled browser speech synthesis), api.openai.com, an empty key, model tts-1 and voice alloy, and no compose change could reach an already-booted volume afterwards. The STT side of this exact reconcile existed; the TTS side had never been written. - owui-patches/hive_rag_env_config.py reconciles audio.tts.engine, model, voice, base URL and API key from the environment, mirroring the STT block, including the paired base-URL-without-key refusal and secret-free logging. - docker-compose.yml sets the five AUDIO_TTS_* variables (OWUI_TTS_ALIAS and OWUI_TTS_VOICE knobs, defaults hive-tts / autumn) following the STT vars. - edge-api serves GET /v1/audio/voices with the provider's real roster (autumn diana hannah austin daniel troy). Open WebUI's get_available_voices fetches that endpoint unauthenticated on a non-OpenAI base URL and falls back to its hardcoded alloy-style list when it fails, which is how the UI kept offering a voice hive-tts rejects (#996). - scripts/test_owui_rag_env_config.py pins all of the above; new Go unit tests cover the voices handler. deploy/litellm/config.yaml verified, no change needed: route-groq-tts exists with api_key os.environ/GROQ_API_KEY and a literal groq/orpheus model id, matching exactly how route-groq-stt is handled. Proof: docs/proof/tts-readaloud/capture.log. Live round trip through LiteLLM route-groq-tts passed via TestLiveVoiceRoundTripThroughLiteLLM (non-empty, non-silent WAV, verbatim STT round trip); direct request returned 200 with 149830 bytes. UI playback capture pending-live-capture post-merge. Fixes #997. Partially addresses #996 by removing alloy from Hive's own surfaces; the 400-to-500 mapping itself remains tracked there.
|
Warning Review limit reached
Next review available in: 2 minutes Limit details: You’ve used the included review currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. How can I continue?Wait for the limit to reset, then comment An organization admin can change what happens after included review limits in Billing. How do review limits work?CodeRabbit enforces per-developer PR review limits within each organization. For paid Pro and Pro+ reviews, CodeRabbit uses a developer's included PR review attempts over the past 7 days to set the current hourly allowance. At typical activity levels, the full plan allowance applies. Higher sustained activity can lower the allowance until earlier attempts leave the 7-day window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: ⛔ Files ignored due to path filters (1)
📒 Files selected for processing (6)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
@coderabbitai review |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
❌ Action failedReview failed.
|
|
|
@coderabbitai review |
|
|
@coderabbitai review |
|
…ve-default and hive-auto (#1099) Owner directive, 2026-08-23: route both `hive-default` and `hive-auto` to OpenRouter free models, and charge 50 percent of what is being charged now. One migration plus one LiteLLM config edit. No Go change on the serving path. ## First, a correction to the premise this task was handed with The brief said `hive-auto` bills at actual upstream cost after PR #1012, and therefore had to be converted back to a fixed price with the "today" figure recovered from ledger rows. That is not what #1012 did. It created a **new** alias, `openrouter-auto`, on a new route `route-openrouter-auto-beta`, in `pricing_mode = 'upstream_actual'`. Its own migration says so out loud: "litellm_model_name is deliberately NOT 'route-openrouter-auto': that name is the retired route id of the pre-existing hive-auto alias, which resolves to a completely different model." `hive-auto` was never touched by it. Both aliases are plain `fixed` rows and have been since the 2026-08-22 restructure, so there is no pricing-mode conversion here and no need to reconstruct a price from ledger magnitudes. The old figures come from the statement that set them. ## Before and after Source for every old rate: `supabase/migrations/20260822_02_catalog_alias_restructure.sql` step 7. | alias | unit | old rate | new rate | source of old rate | |---|---|---|---|---| | hive-default | credits per million **prompt** tokens | 10500 | **5250** | 20260822_02 step 7 | | hive-default | credits per million **completion** tokens | 42000 | **21000** | 20260822_02 step 7 | | hive-default | credits per million cache-read tokens (published, never billed) | 0 | 0 | 20260822_02 step 7 | | hive-default | credits per million cache-write tokens (published, never billed) | 0 | 0 | 20260822_02 step 7 | | hive-auto | credits per million **prompt** tokens | 21000 | **10500** | 20260822_02 step 7 | | hive-auto | credits per million **completion** tokens | 84000 | **42000** | 20260822_02 step 7 | | hive-auto | credits per million cache-read tokens (published, never billed) | 0 | 0 | 20260822_02 step 7 | | hive-auto | credits per million cache-write tokens (published, never billed) | 0 | 0 | 20260822_02 step 7 | Every figure halves with no remainder, so no rounding rule had to be invented and no float can enter. Both aliases stay `pricing_mode = 'fixed'` and `price_unit = 'tokens'`. Why the cache columns are listed at 0 and not omitted: they are the only other price columns on the row, `precedence.go` never reads them, and both were already zeroed by 20260822_02 when the aliases moved to Groq routes declaring no cache support. Half of zero is zero, so they are asserted rather than changed. Naming them is the point: "every unit type the alias bills today" is answered exhaustively rather than by omission. ## Why there is no Go change `apps/edge-api/internal/metering/precedence.go` is the single implementation of the charge arithmetic (D-031). It builds exactly two `UnitCharge` values per request, prompt tokens at `InputPriceCredits` and completion tokens at `OutputPriceCredits`, sums them, divides once by a million and rounds half up. So halving those two columns halves every token charge exactly, for every request shape, with nothing else moving: - **The credit hold does not move**, and should not. For a fixed-price alias `ReservationCredits` returns the flat per-endpoint default (10000 for chat); it only raises a hold above that for an `upstream_actual` alias. The hold is an authorization, released in full at settlement. - **The fail-closed path halves too.** A missing usage block synthesises a token count from response bytes and prices it at the same catalog rate, so it lands at half the old figure rather than at zero or at the old rate. - **Non-token modalities are untouched because they never read these columns.** `/v1/images/*` reserves a hardcoded 5000 credits and `internal/audio` charges flat literals, both alias-independent (D-033, issue #627). See the residuals. - **Which token classes are billed does not change.** Only the rate does. There is a guard for this specifically, because a brief of this shape previously produced a 262x overcharge by widening what was billed. ## The money proof: same request shapes, before and after Produced by calling the production settlement function `inference.CreditsForTokens` directly, once with the old rates and once with the new ones, in the repo's own toolchain container. Not a reimplementation: that is the function both the streaming and the sync settlement paths call. Two figures per shape. **Exact** is quantity times rate, summed, before the single division and round-half-up, which is where "exactly half" is provable. **Settled** is the whole-credit charge the ledger records, and both sides round independently. | alias | request shape | prompt | completion | old exact (credit-millionths) | new exact | exact ratio | old settled | new settled | |---|---|---|---|---|---|---|---|---| | hive-default | typical chat turn | 1200 | 400 | 29400000 | 14700000 | **0.500000** | 29 | 15 | | hive-default | long context read | 32000 | 800 | 369600000 | 184800000 | **0.500000** | 370 | 185 | | hive-default | prompt only | 1500 | 0 | 15750000 | 7875000 | **0.500000** | 16 | 8 | | hive-default | completion only | 0 | 900 | 37800000 | 18900000 | **0.500000** | 38 | 19 | | hive-default | tool-call-only turn | 850 | 40 | 10605000 | 5302500 | **0.500000** | 11 | 5 | | hive-default | fail-closed byte estimate | 72 | 1000 | 42756000 | 21378000 | **0.500000** | 43 | 21 | | hive-default | one token each (floor) | 1 | 1 | 52500 | 26250 | **0.500000** | 1 | 1 | | hive-default | 24 completion tokens (floor boundary) | 0 | 24 | 1008000 | 504000 | **0.500000** | 1 | 1 | | hive-default | negative counts clamped | -5 | -5 | 0 | 0 | n/a | 0 | 0 | | hive-auto | typical chat turn | 1200 | 400 | 58800000 | 29400000 | **0.500000** | 59 | 29 | | hive-auto | long context read | 32000 | 800 | 739200000 | 369600000 | **0.500000** | 739 | 370 | | hive-auto | prompt only | 1500 | 0 | 31500000 | 15750000 | **0.500000** | 32 | 16 | | hive-auto | completion only | 0 | 900 | 75600000 | 37800000 | **0.500000** | 76 | 38 | | hive-auto | tool-call-only turn | 850 | 40 | 21210000 | 10605000 | **0.500000** | 21 | 11 | | hive-auto | fail-closed byte estimate | 72 | 1000 | 85512000 | 42756000 | **0.500000** | 86 | 43 | | hive-auto | one token each (floor) | 1 | 1 | 105000 | 52500 | **0.500000** | 1 | 1 | | hive-auto | 24 completion tokens (floor boundary) | 0 | 24 | 2016000 | 1008000 | **0.500000** | 2 | 1 | | hive-auto | negative counts clamped | -5 | -5 | 0 | 0 | n/a | 0 | 0 | Exactly half on every row, in integer rational arithmetic. Not "roughly half", not a fraction of a percent, not unchanged, and not zero. Stated rather than smoothed over: **the settled whole-credit charge can sit one credit above exactly half**, because the old and new charges each round half up independently from exact rationals in a 2 to 1 ratio (29.4 rounds to 29 while 14.7 rounds to 15). The deviation is bounded at one credit, which is 0.00001 USD at the repo's 100000-credits-per-USD constant. The 1-credit floor in `CreditsForTokens` is the other bounded exception: it is what keeps a sub-credit request off zero, it predates this change, and it is why the two floor rows read 1 and 1. Both are visible in the table rather than hidden behind a "50 percent" claim the integers do not literally satisfy. The replay enforced its own claims: it failed the run if any exact ratio was not exactly one half, if any new charge was zero where the old was positive, or if any new charge exceeded half by more than one credit. Full capture, including the commands: `docs/proof/free-route-aliases-half-price-2026-08-23/README.md`. ## The test that fails on the old rate `apps/control-plane/internal/routing/free_alias_pricing_test.go`, offline, reusing the SQL parser already in that package. Ten guards, all positional: the migration declares its arithmetic in a `-- HALVE| alias | field | old | new` table, the test checks that table is arithmetically half, that its OLD column matches the rates actually in force (pinned here so a later edit to 20260822_02 cannot move the baseline), and then that the value each `UPDATE` **assigns** equals half. A correct comment above a wrong `UPDATE` is caught. Mutation tested rather than asserted to work. Six mutations, six kills, control green: | mutation | result | |---|---| | M1 revert hive-default input to the old 10500 | FAIL `TestFreeAliasPricesAreExactlyHalfTheOldRates` | | M2 hive-default output 21000 becomes hive-auto's 10500 | FAIL `TestFreeAliasPricesAreExactlyHalfTheOldRates` | | M3 route-free-auto loses `supports_batch` and both image flags | FAIL `TestFreeRouteAutoCarriesTheSoleCapabilityFlagsForward` | | M4 provider_model loses the `:free` suffix | FAIL `TestFreeAliasRoutesTargetAFreeOpenRouterModel` | | M5 route-free-default loses `tools_supported` | FAIL `TestFreeRoutesKeepTheCapabilitiesTheirAliasesServeToday` | | M6 the retired Groq routes left enabled | FAIL `TestRetiredGroqRoutesAreDisabledAndRepointed` | M1 is the guard the brief asked for: red on the old rate, green on the new one. The existing `TestCatalogAliasPricesMatchProviderRates` cannot cover these two prices, and that is structural rather than an oversight. It derives every credit figure as `usd_per_million * 1.4 * 100000` (D-032), and its `parseRate` refuses a zero rate outright: "that is a mispricing, not a rate". A free upstream costs zero, so the margin formula yields zero, and zero is refused by `SelectRoute` as unpriceable. **These two prices are owner-set, not cost-derived**, and the halving relation is what replaces the formula as the checkable invariant. ## Routing: which free model, chosen from live data `https://openrouter.ai/api/v1/models`, fetched 2026-08-23: 422 models, 22 priced at zero on both prompt and completion. A literal `openrouter/free` id **does** exist; it is a router, not a model. The capability bar comes from what these aliases serve today. Both current routes declare `tools_supported = true`, which is the column PR #206 routes `tools`, `tool_choice` and `response_format` on, plus `supports_streaming` and `supports_reasoning`. Of the 22 free models, five support all of tools, tool_choice, response_format and structured outputs. Joined to `https://openrouter.ai/api/frontend/v1/all-providers`: | free model | provider | trains on prompts | retention | |---|---|---|---| | dots-studio/dots-3-note-preview:free | AtlasCloud | no | retains, period not published | | z-ai/glm-5.2:free | Decart | no | zero retention | | nvidia/nemotron-3-super-120b-a12b:free | NVIDIA | **YES** | retains | | nvidia/nemotron-nano-9b-v2:free | NVIDIA | **YES** | retains | | liquid/lfm-2.5-2.6b:free | Liquid | **YES** | retains | A false green I shipped and then caught, recorded so the next person does not repeat it: the endpoints API reports NVIDIA's provider as `Nvidia` while the provider directory keys it under displayName `NVIDIA`. Joining on displayName alone misses the record and reports every NVIDIA free endpoint as no-training and zero-retention, the exact opposite of the truth. Join on both fields. Live probes with the project's real key, and this is where the paper ranking fell apart: - **`z-ai/glm-5.2:free`**, the only zero-retention candidate and the strongest model on every published axis: **429 on four of four attempts**, `provider_error_code: upstream_429`, `limit_source: upstream_provider_shared_pool`. That is Decart's shared pool, not our account's limit, and buying credits cannot raise it. Rejected on live evidence rather than on paper. - **`openrouter/free` with no provider preference**: five of five succeeded, but four landed on NVIDIA endpoints and **two of five landed on `nvidia/nemotron-3.5-content-safety:free`, a moderation classifier**, which answered a plain chat prompt with `User Safety: safe` and then with an empty string. Unfit for a chat alias on output quality alone, before the training question. This is also the exact shape issue #689 called a bug. - **`openrouter/free` with `provider: {data_collection: "deny"}`**: five of five, only Cohere and AtlasCloud, tool calls on five of five. The deny filter empirically excludes every NVIDIA endpoint and the classifier. - **`provider: {zdr: true}`**, the stricter form: **404, "No endpoints found matching your data policy (Zero data retention)"**. There is no zero-data-retention free endpoint reachable at all today. - **`dots-studio/dots-3-note-preview:free`**: 200 on a sync completion, on a tools request (a real `tool_calls` with correct arguments and `finish_reason: tool_calls`), on `response_format: {"type":"json_object"}` (valid JSON back) and on a streamed request. `cost: 0` on every response. **Chosen: `dots-studio/dots-3-note-preview:free` for both aliases**, pinned, with `extra_body.provider.data_collection: deny` and `allow_fallbacks: false`. It is the only free model that is simultaneously full parity, live verified working, and served by a provider that does not train on prompts. The `deny` preference is kept even though the model resolves to one provider today: if AtlasCloud's policy changes or a second provider appears, the request fails instead of quietly moving customer prompts to a provider that stores them. ### Capability parity verdict | capability | before (Groq gpt-oss) | after (free model) | verdict | |---|---|---|---| | tools, tool_choice | declared, supported | declared, **live verified** with a real tool call | parity | | response_format, structured outputs | routed on `tools_supported` | supported, **live verified** returning valid JSON | parity | | streaming | declared | declared, **live verified**, 20 chunks | parity | | reasoning / reasoning_effort | declared | declared, model enables reasoning by default | parity | | chat completions, completions, responses | declared | declared | parity | | embeddings | false | false | unchanged | | cache read, cache write | false | false | unchanged, no cache rate published | | batch, image generation, image edit | on route-groq-auto only | carried onto route-free-auto | preserved, see below | | vision (image input) | none reachable | upstream is `text+image->text` | **not declared**, see below | Two notes on that table. `provider_capabilities` has no vision column, so the free model's image-input support is an undeclared property of the upstream rather than a new product claim; the 2026-08-22 note that no customer-reachable vision path exists is now inaccurate at the upstream level and the config comment says so. And the three media flags are a documented status-quo fiction inherited through two migrations: neither gpt-4.1-mini, nor gpt-oss-120b, nor this model generates images. They are carried because `route-groq-auto` is their **sole carrier in the catalog**, `SelectRoute` hard-filters on each flag, and `batchstore` sends `NeedBatch = true` for every batch, so disabling it without handing them on would leave `/v1/batches`, `/v1/images/generations` and `/v1/images/edits` with zero eligible routes **for every alias in the system**. M3 above is the guard. Under-claiming is not a safe default here: both aliases are `pinned` to exactly one route, so `matchesRequestedCapabilities` drops the only candidate, `SelectRoute` returns `ErrRouteNotEligible` and `writeRoutingError` maps it to 422. On a pinned alias an under-claim is a failed request, not a withheld feature. ### Rate limits, and what a user sees at the cap Documented (`https://openrouter.ai/docs/api_reference/limits.md`, whose MDX constants resolve to real numbers): free variants are capped at **20 requests per minute**, and at **50 per day** below 10 dollars of lifetime credit purchases or **1000 per day** at or above it. `GET /api/v1/key` reports `is_free_tier: false` for this account and project memory records a 10 dollar purchase, so 1000 per day is the expected tier. The key endpoint does not expose the purchased total, so that is an expectation, not a measurement; its own `rate_limit` field is documented as deprecated and returns `requests: -1`. Separately and more sharply, the Decart 429 shows a per-provider shared pool can refuse everything regardless of our standing. **What a customer sees at the cap today: a roughly 60 second wait and then a 502 whose message is `context canceled`.** That is issue #1089, filed today and unchanged by this work. `num_retries: 3` with `request_timeout: 45` retries a rate-limited deployment three times, the SDK suites time out at 60 seconds, and the provider's own 429 with its retry hint never leaves the LiteLLM container log. Not fixed here, and the reason is containment rather than appetite. The candidate fix (`router_settings.retry_policy.RateLimitErrorRetries: 0`) belongs in the `litellm_settings` block, which the config sync **preserves verbatim** on a live volume. A file edit there is inert on the box and cannot be verified from a developer machine with no SSH to it. #1089 reaches the same conclusion about itself and asks for a real reproduction. This change does make it materially more likely, since 20 requests per minute is much tighter than Groq's ceiling, so its priority moves from latent to likely. Mitigation that does exist: `hive-small` and `hive-medium` stay on Groq and are now the only customer-reachable Groq chat routes besides the deprecated `hive-fast`. A rate-limited customer has a working alternative, but has to select it; nothing routes them there, because the gateway chat fallbacks were removed deliberately in 20260822_02 and re-adding one is the same inert-on-a-live-volume problem. ### Logging and training terms - **OpenRouter itself** stores no prompts or completions unless the account opts in, and both opt-ins (private input/output logging, and letting OpenRouter use inputs/outputs for a 1 percent discount) are off by default. Request metadata (token counts, latency) is always retained. - **AtlasCloud**, serving the chosen model: `training: false`, `retainsPrompts: true`, no retention period published. - The route carries `provider.data_collection: deny` as a fail-closed guard. - **The contradiction, written down because this product sells data sovereignty:** a strict zero-retention posture is not achievable on any free OpenRouter endpoint today, proved by the `zdr: true` 404. The best available free posture is a provider that does not train but does retain for an unpublished period. The owner directed the move with this on the record. ## Why new route ids rather than repointing the two Groq rows The same reason 20260822_02 gave, and it applies in this direction too. The LiteLLM config sync merges **field by field**: the database owns only `model`, `api_base` and `api_key`, and every other key already on the entry survives so that hand-tuning sticks (`mergeParams`, issue #707). Retiring the route id makes the merge drop the whole stale entry, because a known `route_id` that is no longer active is deleted rather than updated. The rows are **disabled, not deleted**, so the change is reversible and `SelectRoute` filters them out. `price_class` stays `standard`, identical to the routes being replaced. `budget` would arguably describe a free upstream better, but `price_class` feeds `allow_price_class_widening`, and keeping the same value means this repoint cannot change selection behaviour through a second mechanism. The `api_key` follows automatically from `providers.api_key_env` for the row's provider slug, so `openrouter` resolves to `OPENROUTER_API_KEY` without the migration naming a secret. The doubled `openrouter/` prefix on `provider_model` is correct: LiteLLM strips the leading one as its provider selector, exactly as `openrouter/~deepseek/...` and `openrouter/openrouter/auto-beta` already do. The trailing `:free` is load-bearing, not cosmetic: dropping it selects a **paid** endpoint of the same model, so the alias would charge a halved price against a real out-of-pocket cost and nothing else in the tree would notice. M4 is the guard. ## Verification - `go test ./apps/control-plane/... -count=1 -short`: clean, no FAIL, no panic. - `go test ./apps/edge-api/... -count=1 -short`: clean, no FAIL, no panic. - New guards, verbose: 10 of 10 pass; 6 of 6 mutations kill a guard. - `npm run lint:litellm-config`: PASS. `npm run lint:litellm-routing`: PASS. `npm run lint:proof-tokens`: ok, 137 files scanned. - `deploy/litellm/config.yaml` parses and resolves 14 model entries, with both new routes carrying the intended `extra_body`. **Not proved, and it matters:** that the running gateway serves the new route. The config is volume-seeded and the live change arrives through `POST /internal/litellm/sync` reading `provider_routes`. There is no SSH to the demo box from here and CI is the only remote hands, so the on-box confirmation belongs to the deploy run. `deploy-demo-box.yml`'s "Assert model catalog prices agree with the model LiteLLM will call" step is the check that catches a stale volume, and its predicate (`pricing_mode = 'upstream_actual' OR input_price_credits > 0`) covers both of these rows. No ledger row at the new price exists yet either; that needs a served request after the migration applies. **No screenshot**, because this change alters no UI surface. The console catalog table renders whatever price the API returns, and the two figures it will show are the ones proved above. ## Residual questions for the owner Both are stated rather than silently resolved, per the brief. 1. **`hive-auto` now costs exactly twice `hive-default` for the identical model.** Both resolve to the same free upstream, so the surviving 2x gap buys the customer nothing. The directive fixes each alias at 50 percent of **its own** old price, which is what shipped; equalising them is a separate decision. Options: (a) leave as is, honest to the directive, indefensible to a customer who compares them; (b) equalise `hive-auto` to 5250 and 21000, a further reduction that cannot overcharge anyone; (c) give `hive-auto` a distinct larger free model. **Recommendation: (b).** Option (c) is blocked today: every other full-parity free model is served by a provider that trains on prompts, and the one zero-retention candidate returns 429 on every request. 2. **The image path is not halved.** `/v1/images/generations` and `/v1/images/edits` reserve and settle a hardcoded 5000 credits that never reads the alias price, so an image request through `hive-auto` is billed the same as before. It is also not actually servable, since the upstream is a text model and those capability flags have been a documented fiction since 20260414_01. Options: (a) halve the literal to 2500, which halves it for every other alias too and so exceeds this directive's scope; (b) leave it and close the fiction under issue #627; (c) drop the flags, which deletes three endpoints catalog-wide. **Recommendation: (b).** ## Buglog entry ```json {"id":"bug-2026-08-23-openrouter-provider-name-vs-displayname-join","date":"2026-08-23","title":"Joining OpenRouter endpoint provider_name against all-providers displayName silently reports a training provider as zero-retention","error_message":"A capability-and-data-policy scan of OpenRouter's free models reported every NVIDIA free endpoint as training:false, retainsPrompts:false, when NVIDIA's published policy is training:true, retainsPrompts:true","root_cause":"GET /api/v1/models/{slug}/endpoints reports the provider in its `name` form (`Nvidia`) while GET /api/frontend/v1/all-providers keys the record by `displayName` (`NVIDIA`). 18 of 82 providers have name != displayName. A dictionary keyed on displayName therefore misses the lookup entirely, and a default of an empty policy object reads as no-training and zero-retention rather than as unknown, so the failure presents as the safest possible answer instead of an error. The scan looked authoritative and was inverted on exactly the axis a data-sovereignty product cares about.","fix":"Key the provider lookup on both `name` and `displayName`, and treat a missing policy as UNKNOWN rather than as permissive. Re-ran the scan: only two of five full-capability free models are served by non-training providers, not four. Also verified the conclusion independently against live behaviour: with provider.data_collection deny set, five of five requests routed away from every NVIDIA endpoint, which agrees with the corrected join and contradicts the original one.","tags":["openrouter","data-policy","sovereignty","false-green","join-key-mismatch","provider-catalog"]} ``` --- # Part two: the remaining Groq text routes, at prices deliberately unchanged Second owner directive the same day, folded into this PR because it touches the same catalog rows and the same `deploy/litellm/config.yaml`, so a separate PR would conflict by construction. **Move the Groq TEXT and chat-completion models to OpenRouter free as well, to stop the Groq free-tier allowance being drained.** Shipped as a second migration, `20260823_21_groq_text_routes_to_openrouter_free.sql`, deliberately separate from the repricing one. Splitting them is what makes "no price moves here" a structural property of a file rather than a claim in its header. ## Prices for these three do NOT change, and that is the intended outcome | alias | unit | rate before | rate after | note | |---|---|---|---|---| | hive-small | credits per million prompt / completion tokens | 10500 / 42000 | **10500 / 42000** | unchanged | | hive-medium | credits per million prompt / completion tokens | 21000 / 84000 | **21000 / 84000** | unchanged | | hive-fast | credits per million prompt / completion tokens | 10500 / 42000 | **10500 / 42000** | unchanged, deprecated alias | The owner's 50 percent instruction applies to `hive-default` and `hive-auto` only. Serving a same-priced alias from a free upstream widens margin, and that is intended. Stated explicitly so no later reader mistakes it for an oversight, and enforced two ways rather than promised: - `TestGroqFreeRepointTouchesNoPrice` fails if that migration assigns any of `input_price_credits`, `output_price_credits`, either cache column, `pricing_mode` or `price_unit`. The migration does not write `model_aliases` at all, so the guard holds in its strongest form. - The DB-level check reads the prices back after the whole chain applies, and they are byte for byte what they were. `hive-fast`'s cache columns stay at 1 and 4, the stale OpenRouter-era values 20260822_02 examined and deliberately left alone so a deprecated alias would not be repriced on any axis. Not touched here either, for the same reason. ## Audio is out of scope, and the survey behind that `route-groq-stt` (whisper-large-v3) and `route-groq-tts` (Orpheus) are untouched, and `TestGroqFreeRepointLeavesAudioOnGroq` fails if that migration so much as names them in an executable statement. `GROQ_API_KEY` stays required; what Groq no longer serves is chat. "OpenRouter has no audio" would have been an overstatement, so here is the actual picture across all 422 models: - **Text to speech: nothing usable.** Not one model, free or paid, advertises `supported_voices`. The only entries with audio in their output modality are `google/lyria-3-pro-preview` and `google/lyria-3-clip-preview` (free, MUSIC generation, no voice selection) and `openai/gpt-audio` and `openai/gpt-audio-mini` (PAID, speech-to-speech chat). None is an OpenAI-compatible `/v1/audio/speech` endpoint, which is what `internal/audio` speaks. **Orpheus has no replacement at any price.** - **Speech to text: three free models take audio as chat input** (`thinkingmachines/inkling:free`, `thinkingmachines/inkling-small:free`, `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free`). They are not a transcription endpoint: they are chat-completions models, while `route-groq-stt` is a LiteLLM `mode: audio_transcription` route that edge-api's audio handler forwards multipart audio to. Using one would be a new integration, not a repoint. All three are also served by providers whose published policy is training on prompts, a poor destination for dictated speech specifically. Reported, not acted on, exactly as the directive asked. Bengali voice dictation (PR #1079) keeps working. ## Capability parity, probed live because the parameter lists differ This one genuinely needed checking rather than asserting. `dots-3-note-preview` lists `tools`, `tool_choice`, `response_format`, `structured_outputs`, `reasoning`, `include_reasoning`, `max_tokens`, `temperature` and `top_p`, and does **not** list `reasoning_effort`, `stop`, `frequency_penalty`, `presence_penalty`, `seed`, `top_k` or `logprobs`. The Groq gpt-oss models it replaces accept several of those. So: does an unlisted parameter fail, or is it ignored? Twelve request shapes, live: | request shape | result | |---|---| | baseline | 200 | | reasoning_effort=low | 200 | | reasoning_effort=high | 200 | | stop | 200 | | frequency_penalty | 200 | | presence_penalty | 200 | | seed | 200 | | top_k | 200 | | logprobs + top_logprobs | 200 | | n=2 | 200 | | response_format json_schema, strict | 200 | | reasoning_effort + provider.require_parameters | 200 | **All twelve returned 200 with a correct answer. There is no request shape that works today and fails after the repoint**, which is the regression this check exists to rule out. The `require_parameters` case is the interesting one: OpenRouter considers `reasoning_effort` satisfied by this endpoint even under strict parameter enforcement. Two behavioural differences that are not failures, recorded so nobody files them as new bugs: - `reasoning_effort` is accepted and never rejected, but does not reliably modulate effort: `low` produced more reasoning tokens than `high` in the same run. Requests keep working; the knob stops being meaningful. - `n=2` returns one choice. That is OpenRouter's existing behaviour for a provider that does not implement `n`, identical on the routes being replaced. Capability FLAGS are carried across per route rather than uniformly. `route-free-small` and `route-free-medium` mirror their Groq originals including `supports_reasoning = true`; `route-free-fast` keeps `supports_reasoning = false`, which is status-quo preservation of an under-claim `route-groq-fast` has carried since its original 20260331_02 seed and that 20260822_02 examined and deliberately left alone. Widening it would be safe, but a routing migration is not where a deprecated alias should quietly gain a feature. None of these three is a sole carrier of a media flag, so disabling them removes no endpoint. `TestRepointedGroqTextRoutesKeepTheirCapabilities` asserts both directions: the flags they must have, and that they claim none of the three media flags that belong on `route-free-auto`. ## Concentration risk, stated rather than buried **After this change every customer-reachable chat alias except the two paid DeepSeek ones resolves to ONE model at ONE provider on ONE free endpoint.** Five aliases: `hive-default`, `hive-auto`, `hive-small`, `hive-medium`, `hive-fast`. There are no gateway fallbacks, by deliberate design (20260822_02 removed them because a fallback answers from a model the alias was not priced against). So the failure mode **moves rather than disappearing**. OpenRouter documents free variants at **20 requests per minute**, and 50 or 1000 per day depending on lifetime credit purchases; that per-minute ceiling is tighter than the Groq daily allowance this was meant to escape. Separately, the Decart 429 recorded in part one shows an upstream provider shared pool can refuse everything regardless of our account standing, and buying credits cannot raise it. It also **removes the mitigation part one could still point at.** An hour ago a rate-limited customer could select `hive-small` and land on a different provider. There is no such alias now. **What a demo user sees at the cap is unchanged: a roughly 60 second wait and then a 502 whose message is `context canceled`, never the provider's own 429 with its retry hint.** That is issue #1089. Not fixed here for the containment reason given in part one: the candidate fix lives in the `litellm_settings` block the config sync preserves verbatim, so a file edit is inert on a live box and cannot be verified from a developer machine with no SSH to it. This change raises that issue's priority from latent to likely. **This does not close issue #1088.** That is CI consuming the live demo's provider allowance, a different cause with a different owner, and there is deliberately no `Fixes #1088` line anywhere in this PR. ## Part two verification: the full migration chain on a real Postgres Not a file-parsing claim this time. A throwaway `pgvector/pgvector:pg17` container on its own port, seeded with `.github/ci/test-db-bootstrap.sql` and then every file in `supabase/migrations/` in order, exactly as `.github/workflows/ci.yml` does it. Container removed afterwards. `all migrations applied`, no error. Read back from that database: ``` alias_id | in_credits | out_credits | route_id | provider | provider_model --------------+------------+-------------+--------------------+------------+------------------------------------------------- hive-auto | 10500 | 42000 | route-free-auto | openrouter | openrouter/dots-studio/dots-3-note-preview:free hive-default | 5250 | 21000 | route-free-default | openrouter | openrouter/dots-studio/dots-3-note-preview:free hive-fast | 10500 | 42000 | route-free-fast | openrouter | openrouter/dots-studio/dots-3-note-preview:free hive-medium | 21000 | 84000 | route-free-medium | openrouter | openrouter/dots-studio/dots-3-note-preview:free hive-small | 10500 | 42000 | route-free-small | openrouter | openrouter/dots-studio/dots-3-note-preview:free hive-stt | 0 | 4316667 | route-groq-stt | groq | groq/whisper-large-v3 hive-tts | 0 | 3080000 | route-groq-tts | groq | groq/canopylabs/orpheus-v1-english ``` Also confirmed against that database: - **One enabled route per alias.** The violations query returns exactly one row, `hive-embedding-default` with 3, which is the pre-existing exception the integration suite already records in `pendingMultiRouteAliases`. - **The three batch and image flags have a healthy carrier**, `route-free-auto`. `supports_stt` and `supports_tts` are still on the two healthy Groq audio routes. So `/v1/batches`, `/v1/images/generations`, `/v1/images/edits` and both voice endpoints still find an eligible route. - **`policy_mode` moved on no alias.** `hive-fast` is still `latency`, `hive-small` and `hive-medium` still `pinned`, `hive-default` still `stability`, `hive-auto` still `weighted`. Only the route name inside `fallback_order` changed. - **Idempotent.** Both new migrations re-applied to the same database a second time, both clean, prices identical afterwards. - **Integration suite green** against that database: `internal/routing`, `internal/catalog` and `internal/litellmconfig` all `ok` with `-tags integration`. Running the whole `./apps/control-plane/...` integration suite at once also produced failures in `auditworker`, `marketplace` and `tenants`. Those are package-parallelism collisions on one shared database, not this change: each passes on its own against the same database, and none of them reads any of the four tables this branch touches. ## Two test-file notes a reviewer should see rather than discover - `TestHiveFastIsPinnedToGroqAtCorrectedPrice` is **renamed** to `TestHiveFastIsPinnedToOneRouteAtItsUnchangedPrice` and its provider and provider_model expectations updated, because the route legitimately moved. Its two price assertions (10500 and 42000) are unchanged and are now the DB-level guard that this repoint did not quietly reprice three aliases. Its comment records that these figures are no longer derivable from the upstream's cost, which is zero, and must not be "corrected" to match it. - `TestSelectRouteHiveFastResolvesToGroqAtGroqPrice` in `service_test.go` keeps a name that no longer describes the catalog. It is a stub-driven test of the `SelectRoute` algorithm with synthetic route fixtures; it reads neither the database nor any migration, so it is still a valid algorithm test. Left alone rather than renamed in an unrelated file. Push back if you would rather it were renamed. ## Residual question three, added by part two **There is now no chat alias on a second provider.** With every free-model alias behind one 20-requests-per-minute endpoint and no fallbacks, a single upstream 429 takes the whole chat surface down for that minute. Options: (a) leave it, which is what the directive literally asks for; (b) keep one Groq chat route as a deliberate escape hatch, cheapest candidate being the deprecated `hive-fast`, which costs almost nothing in Groq allowance and preserves a working alias to switch to; (c) fix #1089 first so at least the failure is a fast, correct 429 a client library can back off from. **Recommendation: (c) then (b).** Not done here because #1089's fix is not containable from this environment, per the reasoning above.
sakibsadmanshajib
left a comment
There was a problem hiding this comment.
Adversarial review pass (the stage that had not run yet on this PR). Verdict: READY.
Findings
No blocking findings. Three non-blocking observations, none requiring changes:
-
Low, intentional:
GET /v1/audio/voicesis registered without the hk_ authorizer or the tenant voice gate. Verified this is correct and low risk: Open WebUI's own voices fetch sends no Authorization header at all (vendor/open-webui/backend/open_webui/routers/audio.py,get_available_voices), so gating would silently reinstate its hardcoded alloy fallback (#996), and the handler serves a static six name roster with no database, provider, or billing contact and no request derived data. Exposure is limited to six public voice names. -
Info: the failure mode changes by design. With
audio.tts.engineforced to openai, a gateway TTS outage now surfaces an error instead of degrading to browser speech synthesis. That silent degradation was itself the defect (#997: unmetered OS voice playback), and honest failure matches the repo's loud failure posture. No silent fallback remains. -
Info: environment wins for these five keys on every open-webui restart, so admin UI edits to them are reverted at next restart (compose literals are always present). This is the accepted #722 mechanism, identical to STT. One doc nit: the comment "clears AUDIO_TTS_ENGINE to keep whatever its administrator configured" cannot be done through .env alone because the compose value is a literal, it needs an override file. The same wording already ships on the preexisting STT block, so consistency rather than a defect.
Checked clean
- Contract match against the fork:
get_available_voicesfetches GET{base_url}/audio/voiceswith no auth header and readsdata["voices"]entries keyed by id/name;VoicesHandlerreturns exactly that shape and 405s non-GET. The speech path sendsBearer audio.tts.openai.api_keyto{base_url}/audio/speech, which stays behind the existing hk_ authorizer plus tenant voice gate (untouched). - Alias chain: hive-tts is seeded pinned to route-groq-tts (migration 20260717_02), voice default autumn is a real Orpheus voice, LiteLLM config untouched as described.
- No ServeMux pattern conflicts: exact pattern among exact pattern siblings (
/v1/audio/speech,/transcriptions,/translations). Nothing previously served/v1/audio/voices(it 404ed into the alloy fallback), so the new route only removes a failure path. - Secret handling:
audio.tts.openai.api_keyadded to SECRET_KEYS so startup logs never carry it; the paired credential refusal is tested for unset, empty and whitespace key values, persisted values untouched on refusal. - Tests run locally on this branch:
go test ./apps/edge-api/... -count=1 -shortgreen in the toolchain container (25 packages, audio included), andpython3 scripts/test_owui_rag_env_config.pygreen including the new TTS pins. CI rollup green, zero unresolved threads. - Buglog entry carried in the PR body per protocol.
Note
Visual proof (audible read aloud playback captured on the deployed chat surface) is disclosed in the PR body as pending live capture, which can only exist once the merged deployment does. Keep that follow-up capture attached to this task after deploy per the standing visual proof directive; it is a post merge step, not a code defect.
…eway (#1199) Closes #1192 (part one). Parts two and three below are decisions and findings, deliberately not implemented here. ## What this changes One file: `packages/openai-contract/generated/hive-openapi.yaml`, regenerated by its own generator. No hand edits. The generated spec had been produced from an older support matrix and never regenerated, so it annotated 22 live endpoints `planned_for_launch`. Since #1187 that file is served publicly and unauthenticated at `/api/openapi.yaml` for code generators to consume, which means every integrator who pointed tooling at it was told that `POST /v1/chat/completions`, the gateway's primary endpoint, is merely planned. The same applied to `/v1/embeddings`, `/v1/responses`, and all of `/v1/files/*`, `/v1/batches/*`, `/v1/audio/*`, `/v1/images/*` and `/v1/uploads/*`. Regenerated with `sh packages/openai-contract/scripts/generate-matrix.sh`, which runs `packages/openai-contract/scripts/sync_hive_contract.py` against the committed matrix and the pinned upstream document (`upstream/SPEC_VERSION` records the 2026-03-28 download, so the transform is deterministic and reproducible from the repo alone). The diff is 27 insertions and 23 deletions: the 22 `x-hive-status` flips, plus the `/chat/completions` `x-hive-notes` picking up the Phase 20 conditional tool-support text the matrix already carried. There is no formatting churn, which confirms the local PyYAML matches the version that produced the committed file. Running the generator a second time produces byte-identical output, so it is idempotent. ## Verification Re-running the exact comparison the console performs (`diffSpecAgainstMatrix` in `apps/web-console/lib/api-contract.ts`): | | before | after | |---|---|---| | total disagreements | 89 | 67 | | `status_mismatch` | 22 | **0** | | `missing_from_spec` | 67 | 67 | | `missing_from_matrix` | 0 | 0 | The 67 remaining are the benign ones the issue predicted and are unchanged: 51 are `out_of_scope` and dropped from the spec on purpose by the generator, and 16 are Hive-native endpoints that were never in the upstream OpenAI document (part two below). **The direction guard passes.** `tests/unit/console-docs-contract.test.ts` is green, all 12 tests, both with and without a populated env file. The spec operation count still matches the `x-hive-status` annotation count exactly (97 = 97), so the line scan that reads the spec did not stop matching after regeneration. **The served artifact was verified, not assumed.** `hive-web-console-prod:ci` was built from this branch and run, and `GET /api/openapi.yaml` returned HTTP 200 with `content-type: application/yaml`. The served body hashes to `a1aa0cda...`, byte-identical to the regenerated file in the tree. Parsing the served payload confirms `POST /v1/chat/completions`, `/v1/embeddings`, `/v1/responses`, `/v1/files`, `/v1/batches`, `/v1/audio/speech`, `/v1/images/generations` and `/v1/uploads` all now report `supported_now`. The production image was checked specifically because fixing only the dev image would have shipped a page that 500s in production while every local check stayed green. Container unit-test failures unrelated to this change: `tests/unit/ci-web-e2e-secret-free.test.ts` fails in-container by design (it reads a workflow file the Dockerfile never copies in). A further set of render tests fail in the container and drop from 15 files to 8 once an env file is supplied, so they are environment artifacts rather than code. Only three files in the app read the contract at all (`app/api/openapi.yaml/route.ts`, `app/console/docs/page.tsx`, `tests/unit/console-docs-contract.test.ts`) and none of them are in the failing set. ## Part two, a decision for the owner, not implemented here Sixteen endpoints classified `supported_now` in the matrix have no machine-readable schema anywhere: `/v1/rag/*` (6 operations), `/v1/agent/tasks*` (4), `/v1/artifacts*` (3), `/v1/featuregate`, and the Anthropic-compatible `/v1/messages` and `/v1/messages/count_tokens`. The generator is built purely from OpenAI's upstream document, so nothing Hive added itself can appear in its output by construction. A developer cannot generate a client for any of them. Three ways forward, in ascending cost: 1. **Leave it.** The spec stays a pure OpenAI-compatibility document and Hive's own surface stays undocumented in machine-readable form. Zero work, and the docs page keeps reporting the 16 as a visible disagreement rather than hiding them. 2. **Author them by hand and merge them in the generator.** This is cheaper than it looks, because the pattern already exists in this package and is currently orphaned: `packages/openai-contract/spec/paths/` already contains four hand-authored Hive OpenAPI documents (`spend-alerts.yaml`, `invoices.yaml`, `grants.yaml`, `budgets.yaml`) covering the `/api/v1/*` control-plane surface. Nothing reads them. `sync_hive_contract.py` consumes only `upstream/openapi.yaml` and the matrix, so those four files are authored, committed, reviewed, and consumed by nothing. Option 2 is roughly: write 16 operations in that same style, then teach the generator to merge `spec/paths/*.yaml` into its output. The ongoing cost is keeping hand-written schema in step with Go handlers by review discipline alone. 3. **Generate from the handlers.** Highest fidelity and lowest long-term rot, but the services route with plain `net/http.ServeMux` and carry no schema annotations, so this means adding an annotation layer or adopting a framework. Much the largest change. This is the difference between "an OpenAI-compatible gateway" and "a gateway with its own documented API", so it is a product call rather than an engineering one. While mapping this I also found that **`packages/openai-contract/overlays/hive-support-status.yaml` is a third orphaned artefact**. It is a 148-action OpenAPI Overlay document that sets `x-hive-status` per operation, it is not read by the generator (which takes status straight from the matrix), and it is itself stale in exactly the way the generated spec was: it still declares `/chat/completions` `planned_for_launch`. `.wolf/buglog.jsonl` already records it as "consumed by nothing" as of 2026-07-17. Worth deleting or wiring up, so nobody edits it believing it has an effect. ## Part three, the matrix is itself stale against the code Asked for because the console prints the matrix's own `generated` date. Two problems. **The printed date is wrong.** The matrix declares `generated: 2026-03-28`, but the file was last edited 2026-07-28 (#573), and before that 2026-07-22 (#416) and 2026-07-17 (#352). The field is hand-maintained and was not updated by those edits, so the console currently prints a date roughly four months earlier than the file's real content. **Five route families exist in the code and are absent from the matrix entirely**, four of which shipped after the stated generation date: | route family | service | shipped | PR | |---|---|---|---| | `/v1/agent/schedules`, `/v1/agent/schedules/{id}` | edge-api | 2026-08-23 | #1081 | | `/v1/audio/voices` | edge-api | 2026-08-24 | #1079 | | `/v1/tenants/switch` | control-plane | 2026-05-17 | #140 | | `/v1/admin/credit-grants`, `/v1/admin/credit-grants/{id}` | control-plane | 2026-05-08 | #136 | | `/v1/credit-grants/me` | control-plane | 2026-05-08 | #136 | Per the brief I have not fixed this here, because adding matrix rows changes runtime behaviour (see below), would need its own integration-test additions, and would change this PR's diff again. ### The part that needs a decision quickly The matrix is not documentation. `apps/edge-api/cmd/server/main.go:581` wraps the entire mux in `middleware.UnsupportedEndpointMiddleware(m)`, and `matrix.Lookup` (`apps/edge-api/internal/matrix/types.go:44`) returns `StatusUnknown` for any method and path it cannot match exactly or by template. The middleware answers `StatusUnknown` with a 404, `Unknown endpoint`. Replaying that exact lookup algorithm against the committed matrix, all six operations of the two newest edge-api route families resolve to `StatusUnknown`, and therefore to a 404: ``` GET /v1/audio/voices -> UNKNOWN -> 404 GET /v1/agent/schedules -> UNKNOWN -> 404 POST /v1/agent/schedules -> UNKNOWN -> 404 GET /v1/agent/schedules/{id} -> UNKNOWN -> 404 PUT /v1/agent/schedules/{id} -> UNKNOWN -> 404 DELETE /v1/agent/schedules/{id} -> UNKNOWN -> 404 (control: POST /v1/chat/completions -> supported_now) (control: POST /v1/agent/tasks -> supported_now) ``` This is the same failure recorded in `.wolf/buglog.jsonl` as `matrix-missing-proprietary-endpoints` on 2026-07-17, which was a demo blocker. The guard built in response, `apps/edge-api/internal/middleware/unsupported_integration_test.go`, drives the real middleware against the real committed matrix, but it enumerates only the Waves 2-4 routes. It covers neither `/v1/agent/schedules*` nor `/v1/audio/voices`, so the same class of regression landed again without turning anything red. I have verified this by reading the middleware and `Lookup` and by replaying the algorithm, not against a live stack, so please confirm against a running edge-api before acting. If it holds, scheduled agent tasks (#1081) and the Open WebUI voice roster (#1079) are both dead in any deployment, and the fix is matrix rows plus adding those paths to the integration test's table. ## Why this rotted, and the cheapest guard Nothing regenerates the spec in CI. The only reference to this package in `.github/workflows/ci.yml` is an unrelated lint script, so a matrix edit has never been forced to bring the generated spec with it. The repo already has the right pattern a few lines above, for a different artefact: ```yaml - name: Codegen drift — permissions.generated.ts run: | make gen-permissions git diff --exit-code apps/web-console/lib/control-plane/permissions.generated.ts ``` The same shape would have caught this on the day the matrix changed: ```yaml - name: Codegen drift — hive-openapi.yaml run: | sh packages/openai-contract/scripts/generate-matrix.sh git diff --exit-code packages/openai-contract/generated/hive-openapi.yaml ``` Not added here because `.github/workflows/ci.yml` was explicitly out of scope for this task and other lanes are live in it. Recommended as a small follow-up. ## Two smaller findings **The direction guard is now vacuously true, and its allowance is stale.** `console-docs-contract.test.ts` pins the direction of status mismatches rather than their count, which was the right call while a known-stale direction existed. After this PR there are zero status mismatches, so the loop body never executes. The guard is not dead (any *new* mismatch still enters the loop and fails, including the overstating direction it was written to catch), but the allowed direction it whitelists is exactly the bug this PR just fixed, so that specific regression could return silently. The test's comment also now describes a live defect that no longer exists. Per the brief I have not edited the test, since changing a guard in the same PR that changes what it guards deserves a separate decision. The one-line tightening, if wanted: ```ts expect(disagreements.filter((entry) => entry.kind === "status_mismatch")).toEqual([]); ``` That is strictly louder than what is there now, and because the generator stamps `x-hive-status` directly from the matrix, a mismatch is structurally impossible unless somebody edited the matrix without regenerating. In other words, that assertion is the drift guard, in a place CI already runs. **The generator writes a file the repo deleted.** `sync_hive_contract.py` still renders `docs/support-matrix.md`, but that file was removed from the repo by #315 when planning docs moved to the Obsidian vault, and it is not in `.gitignore`. Every run therefore drops a 21 KB untracked file into the tree, which a later `git add -A` would sweep back in and silently restore a deliberately deleted document. I removed it from my tree rather than committing it. Either drop `render_markdown` from the generator or ignore its output, but somebody should pick one. ## Buglog entry To be appended to `.wolf/buglog.jsonl` on `main` in a separate buglog-only PR after this merges, per `.claude/rules/openwolf.md`. ```json {"id":"openapi-spec-stale-understates-gateway","priority":"P1","date":"2026-08-25","error_message":"Publicly served /api/openapi.yaml annotated 22 live endpoints planned_for_launch, including POST /v1/chat/completions, telling every integrator the gateway's primary endpoint was not shipped","root_cause":"packages/openai-contract/generated/hive-openapi.yaml is generated from matrix/support-matrix.json by scripts/sync_hive_contract.py, but nothing in CI regenerates it or fails on drift. The matrix was edited three times (#352, #416, #573) without the generated spec being regenerated, so the spec kept the statuses of a much older matrix. Invisible until #1187 started serving the file publicly and rendering a spec-vs-matrix diff on /console/docs, which surfaced 89 disagreements.","fix":"Ran packages/openai-contract/scripts/generate-matrix.sh to regenerate the spec from the committed matrix and the pinned upstream document, no hand edits. Status mismatches went 22 to 0 and total disagreements 89 to 67 (the remaining 67 are out_of_scope drops and Hive-native endpoints absent from upstream, both by design). Verified the regenerated bytes reach the served artefact by building hive-web-console-prod:ci and hashing the GET /api/openapi.yaml response body. Follow-up recommended: a CI codegen-drift step mirroring the existing permissions.generated.ts one.","tags":["contract","openapi","codegen-drift","docs","ci-gap"]} ``` Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
… guard from the mux (#1203) ## Summary Live authenticated probe against hive-demo-cf confirmed the claim in the dispatch: `GET /v1/audio/voices` and every `/v1/agent/schedules` operation return 404 `{"code":"unknown_endpoint"}` in production, even with a real Supabase session, because they were never added to `packages/openai-contract/matrix/support-matrix.json`. `UnsupportedEndpointMiddleware` wraps the whole `/v1/` mux and answers `StatusUnknown` with 404 before the request ever reaches the feature gate or the handler, so both shipped features (#1079, #1081) have been dead on the live box since they merged. Reading the two handlers turned up two more unlisted operations in the same family while auditing: `GET /v1/agent/tasks/{task_id}/events` and `.../files`, real routes registered on the mux with no matrix entry at all. All eight are added with `supported_now`. ### How I verified (before writing any code) 1. Read `UnsupportedEndpointMiddleware` and `matrix.Lookup`: confirmed the mechanism (default case on `StatusUnknown` returns 404) and that auth (`authSelectorMiddleware`) wraps *outside* the middleware, so an unauthenticated probe cannot distinguish "unlisted" from "unauthorized" — which is exactly why the earlier unauthenticated probe was inconclusive. 2. Minted a real session via the admin one-time-token flow (magic-link mint, no password touched) against the box's self-hosted Supabase, from inside the `control-plane` container (it's the only container holding `SUPABASE_SERVICE_ROLE_KEY`), for `e2e-verified@scubed.com.bd`. 3. Sent raw authenticated HTTP requests to `edge-api:8080` from inside the docker network on the box. `GET /v1/models` (control) returned 200. `GET /v1/audio/voices` and `GET /v1/agent/schedules` both returned: ``` HTTP/1.1 404 Not Found {"error":{"message":"Unknown endpoint: GET /v1/audio/voices","type":"invalid_request_error","param":null,"code":"unknown_endpoint"}} ``` This is the literal `default:` branch body from `UnsupportedEndpointMiddleware`, i.e. `matrix.Lookup` returned `StatusUnknown`. Confirmed, not inferred. 4. Confirmed both endpoints are genuinely registered on the mux (`grep mux.Handle`), so this is purely a matrix data gap, not a routing gap. ## What changed **Data:** 8 new `supported_now` entries in `support-matrix.json`: `GET /v1/audio/voices`, the 5 `/v1/agent/schedules` operations, and `GET /v1/agent/tasks/{task_id}/events` / `.../files`. **Guard, two layers (this is a recurrence — buglog entry `matrix-missing-proprietary-endpoints`, 2026-07-17, was the same class and the guard built then only covers a hand-typed case list nobody extended for these two new families):** - Kept and extended `unsupported_integration_test.go`'s hand-list test with the 8 new cases. This is the only mechanism that can see a new suffix inside a handler's own internal path dispatch (`routeItem`/`routeTaskByID`-style switch) — a raw mux pattern alone cannot. - Added a mux-derived guard that needs no such list. `route_recorder.go` wraps the real `*http.ServeMux`, recording every pattern registered through it. `main()` now refuses to start (`log.Fatal`) if any `/v1/` pattern it actually registered has zero `support-matrix.json` coverage (`assertMatrixCoverage`), and `route_matrix_guard_test.go` exercises the identical check in CI by calling the real registration functions (`registerRAGRoutes`, `registerAgentTaskRoutes`, `registerAgentScheduleRoutes`, `registerInfraRoutes`, `registerMediaFileBatchRoutes`, the newly-extracted `registerAudioVoicesRoute`, and `artifacts.Handler.Register`) with lightweight fakes, mirroring the existing `gated_routes_test.go` pattern of building a real mux from real registration code. - Verified the new guard actually would have caught this bug: replayed it against the pre-fix matrix via `HIVE_MATRIX_PATH_FOR_TEST` and it fails with exactly `/v1/audio/voices, /v1/agent/schedules, /v1/agent/schedules/`. `registerInfraRoutes`, `registerMediaFileBatchRoutes`, the three `gated_routes.go` functions, and `artifacts.Handler.Register` now accept a small local interface (`httpMux` / `muxHandleFunc`) instead of the concrete `*http.ServeMux`, so the recorder can be passed to them with zero behavior change. Existing tests that build a plain `http.NewServeMux()` (`gated_routes_test.go`, `artifacts/handler_test.go`) compile unchanged — `*http.ServeMux` already satisfies both interfaces structurally. **Stated, known limit:** the mux-derived guard cannot see past a registered subtree prefix into a handler's own internal method/suffix switch — that's why the hand-list test stays rather than being deleted. Closing that fully would mean rewriting these proprietary handlers onto Go 1.22+ method+wildcard mux patterns (`mux.HandleFunc("GET /v1/agent/schedules/{id}", ...)`) instead of one prefix registration plus manual dispatch inside. Out of scope here; noted for anyone picking this up further. ## Not fixed here (noted per dispatch instructions) - Nothing in CI regenerates the spec/matrix from source, which is the root cause of the drift. The repo already has the right pattern for this in the `permissions.generated.ts` drift step; applying that pattern to the matrix is a separate change. - `packages/openai-contract/scripts/sync_hive_contract.py` still writes `docs/support-matrix.md`, deleted by #315 and not gitignored, so every run drops an untracked file that a later `git add -A` would silently restore. Also separate. ## Deploy status Deploys to the box are currently blocked by an unrelated failing migration on another branch. This fix will not reach production until that clears — merging this PR alone does not fix the live 404s. ## Test plan - [x] `go build ./apps/edge-api/...` - [x] `gofmt -l` clean on every changed/new file - [x] `go vet ./apps/edge-api/...` clean - [x] `go test ./apps/edge-api/... -count=1` — all packages green, including the extended hand-list test and the new guard test - [x] New guard test replayed against the pre-fix matrix via `HIVE_MATRIX_PATH_FOR_TEST` — fails on exactly the two originally-reported routes, confirming it would have caught this - [x] Live authenticated probe against hive-demo-cf (see Summary) — this is what confirmed the bug in the first place 🤖 Generated with [Claude Code](https://claude.com/claude-code) ## Buglog entry Per `.claude/rules/openwolf.md`, carried here for the follow-up buglog-only PR (not appended to `.wolf/buglog.jsonl` on this branch): ```json {"date":"2026-08-25","tags":["matrix","edge-api","routing","recurrence"],"error_message":"GET /v1/audio/voices and the whole /v1/agent/schedules family returned 404 unknown_endpoint on the live box despite being fully implemented and registered on the mux","root_cause":"UnsupportedEndpointMiddleware 404s any /v1/ path with no support-matrix.json entry, checked before auth/gate/handler; the two route families shipped (#1079, #1081) without matrix entries, and the existing regression guard (unsupported_integration_test.go, added for the same defect on 2026-07-17) only covered a hand-typed case list that nobody extended for these","fix":"Added the 8 missing matrix entries (including two more found while auditing: GET /v1/agent/tasks/{task_id}/events and .../files); added a mux-derived boot-time+CI guard (route_recorder.go, assertMatrixCoverage) that fails on any /v1/ pattern the mux actually registers with zero matrix coverage, so a new route can no longer ship unlisted without a human remembering to update a list"} ``` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for audio voice listing and agent task event, file, and schedule endpoints. * Added startup validation to detect API routes missing from the support matrix. * **Bug Fixes** * Improved route coverage checks to recognize exact paths, normalized paths, and route subtrees. * **Tests** * Added coverage checks for registered routes, including validation that unsupported routes are rejected. * Expanded regression coverage for agent and audio voice routes. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
From the Go review. The exemption itself is unchanged. The constant had been inserted between registerAudioVoicesRoute's doc comment and the function, so godoc attached the issue #1079 paragraph to the constant and left the function undocumented. The constant now sits above that comment with a blank line between them. The bigger point was two tests each naming themselves the complete exemption set. TestOnlyTheDescriptorListIsExemptFromAuth was true when PR #1730 wrote it and became false the moment this change exempted a second route, and false in the direction that still passes, since its table simply does not mention the voice roster. Rather than leave a stale claim next to a fresh one, both move into one table, TestAuthSelectorExemptions, which carries both exempt routes and every negative case for both. The web tools file keeps its reachability test and a note saying where the other half went and why it moved. Three copies of the same middleware closure setup collapse into one helper, exerciseAuthSelector, and the cases run under t.Run so a failure names itself. The real-roster test drops the Name field it decoded and never asserted, and its comment now says plainly what it does not guard: the roster's contents are pinned in internal/audio/handler_voices_test.go, so a hardcoded fallback list would satisfy this test. It guards against no roster, not against the wrong one. Verified the consolidated table can still go red: with the voice roster removed from the exemption, the roster case and the end-to-end test both fail and every other case stays green. Whole package passes, gofmt and go vet clean. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PGAcTcHd3PdD531LaLqXbw
Fixes #997. Partially addresses #996.
What was broken
Open WebUI's text-to-speech had no working engine configured. Every
audio.tts.*key is persistent config, so a first boot seeded engine="" (Open WebUI's bundled browser speech synthesis), base URL https://api.openai.com/v1, an empty API key, model tts-1 and voice alloy, and no compose change could ever reach an already-booted volume afterwards. The STT side of this exact reconcile existed since the dictation fix; the TTS side of the same module had never been written.The UI also kept offering alloy-style voices because Open WebUI's
get_available_voicesfetches GET {base_url}/audio/voices on a non-OpenAI base URL and falls back to its hardcoded OpenAI fallback list when that fetch fails. The gateway served no such endpoint.What changed
Tests
Functional proof
docs/proof/tts-readaloud/capture.log:
Visual proof
Pending-live-capture, honestly: the read-aloud button producing audible playback needs the deployed stack running this branch's images, which does not exist before merge. Post-merge steps: deploy-demo-box.yml auto-deploys main; sign in at chat-hive.scubed.co, click the speaker button on any assistant message, confirm playback and that Settings > Audio TTS Voice lists autumn, diana, hannah, austin, daniel, troy; capture screenshot into docs/proof/tts-readaloud/. A follow-up commit will add it.
Buglog entry
{"error_message":"POST /v1/audio/speech unreachable from chat: OWUI audio.tts.* persisted defaults pointed at api.openai.com with empty key and voice alloy","root_cause":"TTS half of the persistent-config env reconcile never written; all five audio.tts keys are first-boot-wins persistent config, and edge-api served no /v1/audio/voices so the UI offered alloy-style fallback voices","fix":"reconcile audio.tts.engine/model/voice/base_url/api_key from env mirroring STT; compose sets AUDIO_TTS_* vars; edge-api serves GET /v1/audio/voices with the provider roster","tags":"tts,owui,voice,reconcile,compose"}