Skip to content

fix: wire read-aloud TTS to the gateway instead of a dead default - #1079

Merged
sakibsadmanshajib merged 3 commits into
mainfrom
fix/tts-readaloud-gateway
Aug 24, 2026
Merged

sakibsadmanshajib merged 3 commits into
mainfrom
fix/tts-readaloud-gateway

Conversation

@sakibsadmanshajib

Copy link
Copy Markdown
Owner

Fixes #997. Partially addresses #996.

What was broken

Open WebUI's text-to-speech had no working engine configured. Every audio.tts.* key is persistent config, so a first boot seeded engine="" (Open WebUI's bundled browser speech synthesis), base URL https://api.openai.com/v1, an empty API key, model tts-1 and voice alloy, and no compose change could ever reach an already-booted volume afterwards. The STT side of this exact reconcile existed since the dictation fix; the TTS side of the same module had never been written.

The UI also kept offering alloy-style voices because Open WebUI's get_available_voices fetches GET {base_url}/audio/voices on a non-OpenAI base URL and falls back to its hardcoded OpenAI fallback list when that fetch fails. The gateway served no such endpoint.

What changed

  • deploy/docker/owui-patches/hive_rag_env_config.py: extends the existing env reconcile with a TTS block mirroring STT: audio.tts.engine, model, voice, base URL and API key, including the paired base-URL-without-credential refusal and secret-free startup logging.
  • deploy/docker/docker-compose.yml: sets AUDIO_TTS_ENGINE, AUDIO_TTS_OPENAI_API_BASE_URL, AUDIO_TTS_OPENAI_API_KEY (OWUI_SHIM_KEY), AUDIO_TTS_MODEL (${OWUI_TTS_ALIAS:-hive-tts}) and AUDIO_TTS_VOICE (${OWUI_TTS_VOICE:-autumn}), following the STT vars' style.
  • edge-api serves GET /v1/audio/voices returning the provider's real roster (autumn diana hannah austin daniel troy), registered without credentials or the tenant voice gate by design: Open WebUI's own fetch sends no Authorization header, so gating would silently reinstate the alloy fallback. Non-GET answers 405.
  • deploy/litellm/config.yaml verified, no change: route-groq-tts exists with api_key os.environ/GROQ_API_KEY and a literal groq/orpheus model id, exactly matching how route-groq-stt is handled.

Tests

  • scripts/test_owui_rag_env_config.py (CI via make test-scripts): new tests pin the five compose lines, the TTS paired refusal, unset-leaves-persisted semantics, the /v1/audio/voices registration, and secret-free log_summary coverage.
  • Go unit tests cover the voices handler roster and 405.
  • go vet clean on ./apps/edge-api/...

Functional proof

docs/proof/tts-readaloud/capture.log:

  • TestLiveVoiceRoundTripThroughLiteLLM PASS against a LiteLLM proxy started from this branch's deploy/litellm/config.yaml: Handler -> route-groq-tts -> Groq Orpheus returned non-empty non-silent WAV, STT leg transcribed the sentence back verbatim.
  • Direct request with supported voice autumn: 200, 149830 bytes.
  • Unsupported voice alloy still returns 500 through LiteLLM; that 400-to-500 mapping is POST /v1/audio/speech answers 500 for an unsupported voice, including OpenAI's default alloy #996's own contract fix and stays tracked there.

Visual proof

Pending-live-capture, honestly: the read-aloud button producing audible playback needs the deployed stack running this branch's images, which does not exist before merge. Post-merge steps: deploy-demo-box.yml auto-deploys main; sign in at chat-hive.scubed.co, click the speaker button on any assistant message, confirm playback and that Settings > Audio TTS Voice lists autumn, diana, hannah, austin, daniel, troy; capture screenshot into docs/proof/tts-readaloud/. A follow-up commit will add it.

Buglog entry

{"error_message":"POST /v1/audio/speech unreachable from chat: OWUI audio.tts.* persisted defaults pointed at api.openai.com with empty key and voice alloy","root_cause":"TTS half of the persistent-config env reconcile never written; all five audio.tts keys are first-boot-wins persistent config, and edge-api served no /v1/audio/voices so the UI offered alloy-style fallback voices","fix":"reconcile audio.tts.engine/model/voice/base_url/api_key from env mirroring STT; compose sets AUDIO_TTS_* vars; edge-api serves GET /v1/audio/voices with the provider roster","tags":"tts,owui,voice,reconcile,compose"}

Sakib Sadman Shajib and others added 2 commits August 23, 2026 09:25
Open WebUI's text-to-speech was never wired to the gateway (#997): every
audio.tts.* key is persistent config, so a first boot seeded engine=""
(bundled browser speech synthesis), api.openai.com, an empty key, model
tts-1 and voice alloy, and no compose change could reach an already-booted
volume afterwards. The STT side of this exact reconcile existed; the TTS
side had never been written.

- owui-patches/hive_rag_env_config.py reconciles audio.tts.engine, model,
  voice, base URL and API key from the environment, mirroring the STT block,
  including the paired base-URL-without-key refusal and secret-free logging.
- docker-compose.yml sets the five AUDIO_TTS_* variables (OWUI_TTS_ALIAS and
  OWUI_TTS_VOICE knobs, defaults hive-tts / autumn) following the STT vars.
- edge-api serves GET /v1/audio/voices with the provider's real roster
  (autumn diana hannah austin daniel troy). Open WebUI's get_available_voices
  fetches that endpoint unauthenticated on a non-OpenAI base URL and falls
  back to its hardcoded alloy-style list when it fails, which is how the UI
  kept offering a voice hive-tts rejects (#996).
- scripts/test_owui_rag_env_config.py pins all of the above; new Go unit
  tests cover the voices handler.

deploy/litellm/config.yaml verified, no change needed: route-groq-tts exists
with api_key os.environ/GROQ_API_KEY and a literal groq/orpheus model id,
matching exactly how route-groq-stt is handled.

Proof: docs/proof/tts-readaloud/capture.log. Live round trip through LiteLLM
route-groq-tts passed via TestLiveVoiceRoundTripThroughLiteLLM (non-empty,
non-silent WAV, verbatim STT round trip); direct request returned 200 with
149830 bytes. UI playback capture pending-live-capture post-merge.

Fixes #997. Partially addresses #996 by removing alloy from Hive's own
surfaces; the 400-to-500 mapping itself remains tracked there.
@coderabbitai

coderabbitai Bot commented Aug 23, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

@sakibsadmanshajib, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 2 minutes

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

Wait for the limit to reset, then comment @coderabbitai review or push new commits to the PR.

An organization admin can change what happens after included review limits in Billing.

How do review limits work?

CodeRabbit enforces per-developer PR review limits within each organization.

For paid Pro and Pro+ reviews, CodeRabbit uses a developer's included PR review attempts over the past 7 days to set the current hourly allowance. At typical activity levels, the full plan allowance applies. Higher sustained activity can lower the allowance until earlier attempts leave the 7-day window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 70feb4e8-cea0-49ba-9769-6217430d374a

📥 Commits

Reviewing files that changed from the base of the PR and between a84decd and 7a50937.

⛔ Files ignored due to path filters (1)
  • docs/proof/tts-readaloud/capture.log is excluded by !**/*.log
📒 Files selected for processing (6)
  • apps/edge-api/cmd/server/main.go
  • apps/edge-api/internal/audio/handler_voices.go
  • apps/edge-api/internal/audio/handler_voices_test.go
  • deploy/docker/docker-compose.yml
  • deploy/docker/owui-patches/hive_rag_env_config.py
  • scripts/test_owui_rag_env_config.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sakibsadmanshajib
sakibsadmanshajib marked this pull request as ready for review August 23, 2026 13:30
@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

@coderabbitai review

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@coderabbitai

coderabbitai Bot commented Aug 23, 2026 •

Copy link
Copy Markdown
❌ Action failed

Review failed.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai

coderabbitai Bot commented Aug 23, 2026 •

Copy link
Copy Markdown
⚠️ Action not completed

Review rate limited.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Aug 23, 2026 •

Copy link
Copy Markdown
⚠️ Action not completed

Review rate limited.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Aug 23, 2026 •

Copy link
Copy Markdown
⚠️ Action not completed

Review rate limited.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

sakibsadmanshajib added a commit that referenced this pull request Aug 23, 2026
…ve-default and hive-auto (#1099)

Owner directive, 2026-08-23: route both `hive-default` and `hive-auto`
to OpenRouter free models, and charge 50 percent of what is being
charged now.

One migration plus one LiteLLM config edit. No Go change on the serving
path.

## First, a correction to the premise this task was handed with

The brief said `hive-auto` bills at actual upstream cost after PR #1012,
and therefore had to be converted back to a fixed price with the "today"
figure recovered from ledger rows.

That is not what #1012 did. It created a **new** alias,
`openrouter-auto`, on a new route `route-openrouter-auto-beta`, in
`pricing_mode = 'upstream_actual'`. Its own migration says so out loud:
"litellm_model_name is deliberately NOT 'route-openrouter-auto': that
name is the retired route id of the pre-existing hive-auto alias, which
resolves to a completely different model."

`hive-auto` was never touched by it. Both aliases are plain `fixed` rows
and have been since the 2026-08-22 restructure, so there is no
pricing-mode conversion here and no need to reconstruct a price from
ledger magnitudes. The old figures come from the statement that set
them.

## Before and after

Source for every old rate:
`supabase/migrations/20260822_02_catalog_alias_restructure.sql` step 7.

| alias | unit | old rate | new rate | source of old rate |
|---|---|---|---|---|
| hive-default | credits per million **prompt** tokens | 10500 |
**5250** | 20260822_02 step 7 |
| hive-default | credits per million **completion** tokens | 42000 |
**21000** | 20260822_02 step 7 |
| hive-default | credits per million cache-read tokens (published, never
billed) | 0 | 0 | 20260822_02 step 7 |
| hive-default | credits per million cache-write tokens (published,
never billed) | 0 | 0 | 20260822_02 step 7 |
| hive-auto | credits per million **prompt** tokens | 21000 | **10500**
| 20260822_02 step 7 |
| hive-auto | credits per million **completion** tokens | 84000 |
**42000** | 20260822_02 step 7 |
| hive-auto | credits per million cache-read tokens (published, never
billed) | 0 | 0 | 20260822_02 step 7 |
| hive-auto | credits per million cache-write tokens (published, never
billed) | 0 | 0 | 20260822_02 step 7 |

Every figure halves with no remainder, so no rounding rule had to be
invented and no float can enter. Both aliases stay `pricing_mode =
'fixed'` and `price_unit = 'tokens'`.

Why the cache columns are listed at 0 and not omitted: they are the only
other price columns on the row, `precedence.go` never reads them, and
both were already zeroed by 20260822_02 when the aliases moved to Groq
routes declaring no cache support. Half of zero is zero, so they are
asserted rather than changed. Naming them is the point: "every unit type
the alias bills today" is answered exhaustively rather than by omission.

## Why there is no Go change

`apps/edge-api/internal/metering/precedence.go` is the single
implementation of the charge arithmetic (D-031). It builds exactly two
`UnitCharge` values per request, prompt tokens at `InputPriceCredits`
and completion tokens at `OutputPriceCredits`, sums them, divides once
by a million and rounds half up. So halving those two columns halves
every token charge exactly, for every request shape, with nothing else
moving:

- **The credit hold does not move**, and should not. For a fixed-price
alias `ReservationCredits` returns the flat per-endpoint default (10000
for chat); it only raises a hold above that for an `upstream_actual`
alias. The hold is an authorization, released in full at settlement.
- **The fail-closed path halves too.** A missing usage block synthesises
a token count from response bytes and prices it at the same catalog
rate, so it lands at half the old figure rather than at zero or at the
old rate.
- **Non-token modalities are untouched because they never read these
columns.** `/v1/images/*` reserves a hardcoded 5000 credits and
`internal/audio` charges flat literals, both alias-independent (D-033,
issue #627). See the residuals.
- **Which token classes are billed does not change.** Only the rate
does. There is a guard for this specifically, because a brief of this
shape previously produced a 262x overcharge by widening what was billed.

## The money proof: same request shapes, before and after

Produced by calling the production settlement function
`inference.CreditsForTokens` directly, once with the old rates and once
with the new ones, in the repo's own toolchain container. Not a
reimplementation: that is the function both the streaming and the sync
settlement paths call.

Two figures per shape. **Exact** is quantity times rate, summed, before
the single division and round-half-up, which is where "exactly half" is
provable. **Settled** is the whole-credit charge the ledger records, and
both sides round independently.

| alias | request shape | prompt | completion | old exact
(credit-millionths) | new exact | exact ratio | old settled | new
settled |
|---|---|---|---|---|---|---|---|---|
| hive-default | typical chat turn | 1200 | 400 | 29400000 | 14700000 |
**0.500000** | 29 | 15 |
| hive-default | long context read | 32000 | 800 | 369600000 | 184800000
| **0.500000** | 370 | 185 |
| hive-default | prompt only | 1500 | 0 | 15750000 | 7875000 |
**0.500000** | 16 | 8 |
| hive-default | completion only | 0 | 900 | 37800000 | 18900000 |
**0.500000** | 38 | 19 |
| hive-default | tool-call-only turn | 850 | 40 | 10605000 | 5302500 |
**0.500000** | 11 | 5 |
| hive-default | fail-closed byte estimate | 72 | 1000 | 42756000 |
21378000 | **0.500000** | 43 | 21 |
| hive-default | one token each (floor) | 1 | 1 | 52500 | 26250 |
**0.500000** | 1 | 1 |
| hive-default | 24 completion tokens (floor boundary) | 0 | 24 |
1008000 | 504000 | **0.500000** | 1 | 1 |
| hive-default | negative counts clamped | -5 | -5 | 0 | 0 | n/a | 0 | 0
|
| hive-auto | typical chat turn | 1200 | 400 | 58800000 | 29400000 |
**0.500000** | 59 | 29 |
| hive-auto | long context read | 32000 | 800 | 739200000 | 369600000 |
**0.500000** | 739 | 370 |
| hive-auto | prompt only | 1500 | 0 | 31500000 | 15750000 |
**0.500000** | 32 | 16 |
| hive-auto | completion only | 0 | 900 | 75600000 | 37800000 |
**0.500000** | 76 | 38 |
| hive-auto | tool-call-only turn | 850 | 40 | 21210000 | 10605000 |
**0.500000** | 21 | 11 |
| hive-auto | fail-closed byte estimate | 72 | 1000 | 85512000 |
42756000 | **0.500000** | 86 | 43 |
| hive-auto | one token each (floor) | 1 | 1 | 105000 | 52500 |
**0.500000** | 1 | 1 |
| hive-auto | 24 completion tokens (floor boundary) | 0 | 24 | 2016000 |
1008000 | **0.500000** | 2 | 1 |
| hive-auto | negative counts clamped | -5 | -5 | 0 | 0 | n/a | 0 | 0 |

Exactly half on every row, in integer rational arithmetic. Not "roughly
half", not a fraction of a percent, not unchanged, and not zero.

Stated rather than smoothed over: **the settled whole-credit charge can
sit one credit above exactly half**, because the old and new charges
each round half up independently from exact rationals in a 2 to 1 ratio
(29.4 rounds to 29 while 14.7 rounds to 15). The deviation is bounded at
one credit, which is 0.00001 USD at the repo's 100000-credits-per-USD
constant. The 1-credit floor in `CreditsForTokens` is the other bounded
exception: it is what keeps a sub-credit request off zero, it predates
this change, and it is why the two floor rows read 1 and 1. Both are
visible in the table rather than hidden behind a "50 percent" claim the
integers do not literally satisfy.

The replay enforced its own claims: it failed the run if any exact ratio
was not exactly one half, if any new charge was zero where the old was
positive, or if any new charge exceeded half by more than one credit.

Full capture, including the commands:
`docs/proof/free-route-aliases-half-price-2026-08-23/README.md`.

## The test that fails on the old rate

`apps/control-plane/internal/routing/free_alias_pricing_test.go`,
offline, reusing the SQL parser already in that package. Ten guards, all
positional: the migration declares its arithmetic in a `-- HALVE| alias
| field | old | new` table, the test checks that table is arithmetically
half, that its OLD column matches the rates actually in force (pinned
here so a later edit to 20260822_02 cannot move the baseline), and then
that the value each `UPDATE` **assigns** equals half. A correct comment
above a wrong `UPDATE` is caught.

Mutation tested rather than asserted to work. Six mutations, six kills,
control green:

| mutation | result |
|---|---|
| M1 revert hive-default input to the old 10500 | FAIL
`TestFreeAliasPricesAreExactlyHalfTheOldRates` |
| M2 hive-default output 21000 becomes hive-auto's 10500 | FAIL
`TestFreeAliasPricesAreExactlyHalfTheOldRates` |
| M3 route-free-auto loses `supports_batch` and both image flags | FAIL
`TestFreeRouteAutoCarriesTheSoleCapabilityFlagsForward` |
| M4 provider_model loses the `:free` suffix | FAIL
`TestFreeAliasRoutesTargetAFreeOpenRouterModel` |
| M5 route-free-default loses `tools_supported` | FAIL
`TestFreeRoutesKeepTheCapabilitiesTheirAliasesServeToday` |
| M6 the retired Groq routes left enabled | FAIL
`TestRetiredGroqRoutesAreDisabledAndRepointed` |

M1 is the guard the brief asked for: red on the old rate, green on the
new one.

The existing `TestCatalogAliasPricesMatchProviderRates` cannot cover
these two prices, and that is structural rather than an oversight. It
derives every credit figure as `usd_per_million * 1.4 * 100000` (D-032),
and its `parseRate` refuses a zero rate outright: "that is a mispricing,
not a rate". A free upstream costs zero, so the margin formula yields
zero, and zero is refused by `SelectRoute` as unpriceable. **These two
prices are owner-set, not cost-derived**, and the halving relation is
what replaces the formula as the checkable invariant.

## Routing: which free model, chosen from live data

`https://openrouter.ai/api/v1/models`, fetched 2026-08-23: 422 models,
22 priced at zero on both prompt and completion. A literal
`openrouter/free` id **does** exist; it is a router, not a model.

The capability bar comes from what these aliases serve today. Both
current routes declare `tools_supported = true`, which is the column PR
#206 routes `tools`, `tool_choice` and `response_format` on, plus
`supports_streaming` and `supports_reasoning`. Of the 22 free models,
five support all of tools, tool_choice, response_format and structured
outputs. Joined to
`https://openrouter.ai/api/frontend/v1/all-providers`:

| free model | provider | trains on prompts | retention |
|---|---|---|---|
| dots-studio/dots-3-note-preview:free | AtlasCloud | no | retains,
period not published |
| z-ai/glm-5.2:free | Decart | no | zero retention |
| nvidia/nemotron-3-super-120b-a12b:free | NVIDIA | **YES** | retains |
| nvidia/nemotron-nano-9b-v2:free | NVIDIA | **YES** | retains |
| liquid/lfm-2.5-2.6b:free | Liquid | **YES** | retains |

A false green I shipped and then caught, recorded so the next person
does not repeat it: the endpoints API reports NVIDIA's provider as
`Nvidia` while the provider directory keys it under displayName
`NVIDIA`. Joining on displayName alone misses the record and reports
every NVIDIA free endpoint as no-training and zero-retention, the exact
opposite of the truth. Join on both fields.

Live probes with the project's real key, and this is where the paper
ranking fell apart:

- **`z-ai/glm-5.2:free`**, the only zero-retention candidate and the
strongest model on every published axis: **429 on four of four
attempts**, `provider_error_code: upstream_429`, `limit_source:
upstream_provider_shared_pool`. That is Decart's shared pool, not our
account's limit, and buying credits cannot raise it. Rejected on live
evidence rather than on paper.
- **`openrouter/free` with no provider preference**: five of five
succeeded, but four landed on NVIDIA endpoints and **two of five landed
on `nvidia/nemotron-3.5-content-safety:free`, a moderation classifier**,
which answered a plain chat prompt with `User Safety: safe` and then
with an empty string. Unfit for a chat alias on output quality alone,
before the training question. This is also the exact shape issue #689
called a bug.
- **`openrouter/free` with `provider: {data_collection: "deny"}`**: five
of five, only Cohere and AtlasCloud, tool calls on five of five. The
deny filter empirically excludes every NVIDIA endpoint and the
classifier.
- **`provider: {zdr: true}`**, the stricter form: **404, "No endpoints
found matching your data policy (Zero data retention)"**. There is no
zero-data-retention free endpoint reachable at all today.
- **`dots-studio/dots-3-note-preview:free`**: 200 on a sync completion,
on a tools request (a real `tool_calls` with correct arguments and
`finish_reason: tool_calls`), on `response_format:
{"type":"json_object"}` (valid JSON back) and on a streamed request.
`cost: 0` on every response.

**Chosen: `dots-studio/dots-3-note-preview:free` for both aliases**,
pinned, with `extra_body.provider.data_collection: deny` and
`allow_fallbacks: false`. It is the only free model that is
simultaneously full parity, live verified working, and served by a
provider that does not train on prompts. The `deny` preference is kept
even though the model resolves to one provider today: if AtlasCloud's
policy changes or a second provider appears, the request fails instead
of quietly moving customer prompts to a provider that stores them.

### Capability parity verdict

| capability | before (Groq gpt-oss) | after (free model) | verdict |
|---|---|---|---|
| tools, tool_choice | declared, supported | declared, **live verified**
with a real tool call | parity |
| response_format, structured outputs | routed on `tools_supported` |
supported, **live verified** returning valid JSON | parity |
| streaming | declared | declared, **live verified**, 20 chunks | parity
|
| reasoning / reasoning_effort | declared | declared, model enables
reasoning by default | parity |
| chat completions, completions, responses | declared | declared |
parity |
| embeddings | false | false | unchanged |
| cache read, cache write | false | false | unchanged, no cache rate
published |
| batch, image generation, image edit | on route-groq-auto only |
carried onto route-free-auto | preserved, see below |
| vision (image input) | none reachable | upstream is `text+image->text`
| **not declared**, see below |

Two notes on that table. `provider_capabilities` has no vision column,
so the free model's image-input support is an undeclared property of the
upstream rather than a new product claim; the 2026-08-22 note that no
customer-reachable vision path exists is now inaccurate at the upstream
level and the config comment says so. And the three media flags are a
documented status-quo fiction inherited through two migrations: neither
gpt-4.1-mini, nor gpt-oss-120b, nor this model generates images. They
are carried because `route-groq-auto` is their **sole carrier in the
catalog**, `SelectRoute` hard-filters on each flag, and `batchstore`
sends `NeedBatch = true` for every batch, so disabling it without
handing them on would leave `/v1/batches`, `/v1/images/generations` and
`/v1/images/edits` with zero eligible routes **for every alias in the
system**. M3 above is the guard.

Under-claiming is not a safe default here: both aliases are `pinned` to
exactly one route, so `matchesRequestedCapabilities` drops the only
candidate, `SelectRoute` returns `ErrRouteNotEligible` and
`writeRoutingError` maps it to 422. On a pinned alias an under-claim is
a failed request, not a withheld feature.

### Rate limits, and what a user sees at the cap

Documented (`https://openrouter.ai/docs/api_reference/limits.md`, whose
MDX constants resolve to real numbers): free variants are capped at **20
requests per minute**, and at **50 per day** below 10 dollars of
lifetime credit purchases or **1000 per day** at or above it. `GET
/api/v1/key` reports `is_free_tier: false` for this account and project
memory records a 10 dollar purchase, so 1000 per day is the expected
tier. The key endpoint does not expose the purchased total, so that is
an expectation, not a measurement; its own `rate_limit` field is
documented as deprecated and returns `requests: -1`. Separately and more
sharply, the Decart 429 shows a per-provider shared pool can refuse
everything regardless of our standing.

**What a customer sees at the cap today: a roughly 60 second wait and
then a 502 whose message is `context canceled`.** That is issue #1089,
filed today and unchanged by this work. `num_retries: 3` with
`request_timeout: 45` retries a rate-limited deployment three times, the
SDK suites time out at 60 seconds, and the provider's own 429 with its
retry hint never leaves the LiteLLM container log.

Not fixed here, and the reason is containment rather than appetite. The
candidate fix (`router_settings.retry_policy.RateLimitErrorRetries: 0`)
belongs in the `litellm_settings` block, which the config sync
**preserves verbatim** on a live volume. A file edit there is inert on
the box and cannot be verified from a developer machine with no SSH to
it. #1089 reaches the same conclusion about itself and asks for a real
reproduction. This change does make it materially more likely, since 20
requests per minute is much tighter than Groq's ceiling, so its priority
moves from latent to likely.

Mitigation that does exist: `hive-small` and `hive-medium` stay on Groq
and are now the only customer-reachable Groq chat routes besides the
deprecated `hive-fast`. A rate-limited customer has a working
alternative, but has to select it; nothing routes them there, because
the gateway chat fallbacks were removed deliberately in 20260822_02 and
re-adding one is the same inert-on-a-live-volume problem.

### Logging and training terms

- **OpenRouter itself** stores no prompts or completions unless the
account opts in, and both opt-ins (private input/output logging, and
letting OpenRouter use inputs/outputs for a 1 percent discount) are off
by default. Request metadata (token counts, latency) is always retained.
- **AtlasCloud**, serving the chosen model: `training: false`,
`retainsPrompts: true`, no retention period published.
- The route carries `provider.data_collection: deny` as a fail-closed
guard.
- **The contradiction, written down because this product sells data
sovereignty:** a strict zero-retention posture is not achievable on any
free OpenRouter endpoint today, proved by the `zdr: true` 404. The best
available free posture is a provider that does not train but does retain
for an unpublished period. The owner directed the move with this on the
record.

## Why new route ids rather than repointing the two Groq rows

The same reason 20260822_02 gave, and it applies in this direction too.
The LiteLLM config sync merges **field by field**: the database owns
only `model`, `api_base` and `api_key`, and every other key already on
the entry survives so that hand-tuning sticks (`mergeParams`, issue
#707). Retiring the route id makes the merge drop the whole stale entry,
because a known `route_id` that is no longer active is deleted rather
than updated. The rows are **disabled, not deleted**, so the change is
reversible and `SelectRoute` filters them out.

`price_class` stays `standard`, identical to the routes being replaced.
`budget` would arguably describe a free upstream better, but
`price_class` feeds `allow_price_class_widening`, and keeping the same
value means this repoint cannot change selection behaviour through a
second mechanism.

The `api_key` follows automatically from `providers.api_key_env` for the
row's provider slug, so `openrouter` resolves to `OPENROUTER_API_KEY`
without the migration naming a secret. The doubled `openrouter/` prefix
on `provider_model` is correct: LiteLLM strips the leading one as its
provider selector, exactly as `openrouter/~deepseek/...` and
`openrouter/openrouter/auto-beta` already do. The trailing `:free` is
load-bearing, not cosmetic: dropping it selects a **paid** endpoint of
the same model, so the alias would charge a halved price against a real
out-of-pocket cost and nothing else in the tree would notice. M4 is the
guard.

## Verification

- `go test ./apps/control-plane/... -count=1 -short`: clean, no FAIL, no
panic.
- `go test ./apps/edge-api/... -count=1 -short`: clean, no FAIL, no
panic.
- New guards, verbose: 10 of 10 pass; 6 of 6 mutations kill a guard.
- `npm run lint:litellm-config`: PASS. `npm run lint:litellm-routing`:
PASS. `npm run lint:proof-tokens`: ok, 137 files scanned.
- `deploy/litellm/config.yaml` parses and resolves 14 model entries,
with both new routes carrying the intended `extra_body`.

**Not proved, and it matters:** that the running gateway serves the new
route. The config is volume-seeded and the live change arrives through
`POST /internal/litellm/sync` reading `provider_routes`. There is no SSH
to the demo box from here and CI is the only remote hands, so the on-box
confirmation belongs to the deploy run. `deploy-demo-box.yml`'s "Assert
model catalog prices agree with the model LiteLLM will call" step is the
check that catches a stale volume, and its predicate (`pricing_mode =
'upstream_actual' OR input_price_credits > 0`) covers both of these
rows. No ledger row at the new price exists yet either; that needs a
served request after the migration applies.

**No screenshot**, because this change alters no UI surface. The console
catalog table renders whatever price the API returns, and the two
figures it will show are the ones proved above.

## Residual questions for the owner

Both are stated rather than silently resolved, per the brief.

1. **`hive-auto` now costs exactly twice `hive-default` for the
identical model.** Both resolve to the same free upstream, so the
surviving 2x gap buys the customer nothing. The directive fixes each
alias at 50 percent of **its own** old price, which is what shipped;
equalising them is a separate decision. Options: (a) leave as is, honest
to the directive, indefensible to a customer who compares them; (b)
equalise `hive-auto` to 5250 and 21000, a further reduction that cannot
overcharge anyone; (c) give `hive-auto` a distinct larger free model.
**Recommendation: (b).** Option (c) is blocked today: every other
full-parity free model is served by a provider that trains on prompts,
and the one zero-retention candidate returns 429 on every request.

2. **The image path is not halved.** `/v1/images/generations` and
`/v1/images/edits` reserve and settle a hardcoded 5000 credits that
never reads the alias price, so an image request through `hive-auto` is
billed the same as before. It is also not actually servable, since the
upstream is a text model and those capability flags have been a
documented fiction since 20260414_01. Options: (a) halve the literal to
2500, which halves it for every other alias too and so exceeds this
directive's scope; (b) leave it and close the fiction under issue #627;
(c) drop the flags, which deletes three endpoints catalog-wide.
**Recommendation: (b).**

## Buglog entry

```json
{"id":"bug-2026-08-23-openrouter-provider-name-vs-displayname-join","date":"2026-08-23","title":"Joining OpenRouter endpoint provider_name against all-providers displayName silently reports a training provider as zero-retention","error_message":"A capability-and-data-policy scan of OpenRouter's free models reported every NVIDIA free endpoint as training:false, retainsPrompts:false, when NVIDIA's published policy is training:true, retainsPrompts:true","root_cause":"GET /api/v1/models/{slug}/endpoints reports the provider in its `name` form (`Nvidia`) while GET /api/frontend/v1/all-providers keys the record by `displayName` (`NVIDIA`). 18 of 82 providers have name != displayName. A dictionary keyed on displayName therefore misses the lookup entirely, and a default of an empty policy object reads as no-training and zero-retention rather than as unknown, so the failure presents as the safest possible answer instead of an error. The scan looked authoritative and was inverted on exactly the axis a data-sovereignty product cares about.","fix":"Key the provider lookup on both `name` and `displayName`, and treat a missing policy as UNKNOWN rather than as permissive. Re-ran the scan: only two of five full-capability free models are served by non-training providers, not four. Also verified the conclusion independently against live behaviour: with provider.data_collection deny set, five of five requests routed away from every NVIDIA endpoint, which agrees with the corrected join and contradicts the original one.","tags":["openrouter","data-policy","sovereignty","false-green","join-key-mismatch","provider-catalog"]}
```

---

# Part two: the remaining Groq text routes, at prices deliberately
unchanged

Second owner directive the same day, folded into this PR because it
touches the same catalog rows and the same `deploy/litellm/config.yaml`,
so a separate PR would conflict by construction.

**Move the Groq TEXT and chat-completion models to OpenRouter free as
well, to stop the Groq free-tier allowance being drained.**

Shipped as a second migration,
`20260823_21_groq_text_routes_to_openrouter_free.sql`, deliberately
separate from the repricing one. Splitting them is what makes "no price
moves here" a structural property of a file rather than a claim in its
header.

## Prices for these three do NOT change, and that is the intended
outcome

| alias | unit | rate before | rate after | note |
|---|---|---|---|---|
| hive-small | credits per million prompt / completion tokens | 10500 /
42000 | **10500 / 42000** | unchanged |
| hive-medium | credits per million prompt / completion tokens | 21000 /
84000 | **21000 / 84000** | unchanged |
| hive-fast | credits per million prompt / completion tokens | 10500 /
42000 | **10500 / 42000** | unchanged, deprecated alias |

The owner's 50 percent instruction applies to `hive-default` and
`hive-auto` only. Serving a same-priced alias from a free upstream
widens margin, and that is intended. Stated explicitly so no later
reader mistakes it for an oversight, and enforced two ways rather than
promised:

- `TestGroqFreeRepointTouchesNoPrice` fails if that migration assigns
any of `input_price_credits`, `output_price_credits`, either cache
column, `pricing_mode` or `price_unit`. The migration does not write
`model_aliases` at all, so the guard holds in its strongest form.
- The DB-level check reads the prices back after the whole chain
applies, and they are byte for byte what they were.

`hive-fast`'s cache columns stay at 1 and 4, the stale OpenRouter-era
values 20260822_02 examined and deliberately left alone so a deprecated
alias would not be repriced on any axis. Not touched here either, for
the same reason.

## Audio is out of scope, and the survey behind that

`route-groq-stt` (whisper-large-v3) and `route-groq-tts` (Orpheus) are
untouched, and `TestGroqFreeRepointLeavesAudioOnGroq` fails if that
migration so much as names them in an executable statement.
`GROQ_API_KEY` stays required; what Groq no longer serves is chat.

"OpenRouter has no audio" would have been an overstatement, so here is
the actual picture across all 422 models:

- **Text to speech: nothing usable.** Not one model, free or paid,
advertises `supported_voices`. The only entries with audio in their
output modality are `google/lyria-3-pro-preview` and
`google/lyria-3-clip-preview` (free, MUSIC generation, no voice
selection) and `openai/gpt-audio` and `openai/gpt-audio-mini` (PAID,
speech-to-speech chat). None is an OpenAI-compatible `/v1/audio/speech`
endpoint, which is what `internal/audio` speaks. **Orpheus has no
replacement at any price.**
- **Speech to text: three free models take audio as chat input**
(`thinkingmachines/inkling:free`, `thinkingmachines/inkling-small:free`,
`nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free`). They are not a
transcription endpoint: they are chat-completions models, while
`route-groq-stt` is a LiteLLM `mode: audio_transcription` route that
edge-api's audio handler forwards multipart audio to. Using one would be
a new integration, not a repoint. All three are also served by providers
whose published policy is training on prompts, a poor destination for
dictated speech specifically.

Reported, not acted on, exactly as the directive asked. Bengali voice
dictation (PR #1079) keeps working.

## Capability parity, probed live because the parameter lists differ

This one genuinely needed checking rather than asserting.
`dots-3-note-preview` lists `tools`, `tool_choice`, `response_format`,
`structured_outputs`, `reasoning`, `include_reasoning`, `max_tokens`,
`temperature` and `top_p`, and does **not** list `reasoning_effort`,
`stop`, `frequency_penalty`, `presence_penalty`, `seed`, `top_k` or
`logprobs`. The Groq gpt-oss models it replaces accept several of those.
So: does an unlisted parameter fail, or is it ignored?

Twelve request shapes, live:

| request shape | result |
|---|---|
| baseline | 200 |
| reasoning_effort=low | 200 |
| reasoning_effort=high | 200 |
| stop | 200 |
| frequency_penalty | 200 |
| presence_penalty | 200 |
| seed | 200 |
| top_k | 200 |
| logprobs + top_logprobs | 200 |
| n=2 | 200 |
| response_format json_schema, strict | 200 |
| reasoning_effort + provider.require_parameters | 200 |

**All twelve returned 200 with a correct answer. There is no request
shape that works today and fails after the repoint**, which is the
regression this check exists to rule out. The `require_parameters` case
is the interesting one: OpenRouter considers `reasoning_effort`
satisfied by this endpoint even under strict parameter enforcement.

Two behavioural differences that are not failures, recorded so nobody
files them as new bugs:

- `reasoning_effort` is accepted and never rejected, but does not
reliably modulate effort: `low` produced more reasoning tokens than
`high` in the same run. Requests keep working; the knob stops being
meaningful.
- `n=2` returns one choice. That is OpenRouter's existing behaviour for
a provider that does not implement `n`, identical on the routes being
replaced.

Capability FLAGS are carried across per route rather than uniformly.
`route-free-small` and `route-free-medium` mirror their Groq originals
including `supports_reasoning = true`; `route-free-fast` keeps
`supports_reasoning = false`, which is status-quo preservation of an
under-claim `route-groq-fast` has carried since its original 20260331_02
seed and that 20260822_02 examined and deliberately left alone. Widening
it would be safe, but a routing migration is not where a deprecated
alias should quietly gain a feature.

None of these three is a sole carrier of a media flag, so disabling them
removes no endpoint. `TestRepointedGroqTextRoutesKeepTheirCapabilities`
asserts both directions: the flags they must have, and that they claim
none of the three media flags that belong on `route-free-auto`.

## Concentration risk, stated rather than buried

**After this change every customer-reachable chat alias except the two
paid DeepSeek ones resolves to ONE model at ONE provider on ONE free
endpoint.** Five aliases: `hive-default`, `hive-auto`, `hive-small`,
`hive-medium`, `hive-fast`. There are no gateway fallbacks, by
deliberate design (20260822_02 removed them because a fallback answers
from a model the alias was not priced against).

So the failure mode **moves rather than disappearing**. OpenRouter
documents free variants at **20 requests per minute**, and 50 or 1000
per day depending on lifetime credit purchases; that per-minute ceiling
is tighter than the Groq daily allowance this was meant to escape.
Separately, the Decart 429 recorded in part one shows an upstream
provider shared pool can refuse everything regardless of our account
standing, and buying credits cannot raise it.

It also **removes the mitigation part one could still point at.** An
hour ago a rate-limited customer could select `hive-small` and land on a
different provider. There is no such alias now.

**What a demo user sees at the cap is unchanged: a roughly 60 second
wait and then a 502 whose message is `context canceled`, never the
provider's own 429 with its retry hint.** That is issue #1089. Not fixed
here for the containment reason given in part one: the candidate fix
lives in the `litellm_settings` block the config sync preserves
verbatim, so a file edit is inert on a live box and cannot be verified
from a developer machine with no SSH to it. This change raises that
issue's priority from latent to likely.

**This does not close issue #1088.** That is CI consuming the live
demo's provider allowance, a different cause with a different owner, and
there is deliberately no `Fixes #1088` line anywhere in this PR.

## Part two verification: the full migration chain on a real Postgres

Not a file-parsing claim this time. A throwaway `pgvector/pgvector:pg17`
container on its own port, seeded with
`.github/ci/test-db-bootstrap.sql` and then every file in
`supabase/migrations/` in order, exactly as `.github/workflows/ci.yml`
does it. Container removed afterwards. `all migrations applied`, no
error.

Read back from that database:

```
   alias_id   | in_credits | out_credits |     route_id       |  provider  |                provider_model
--------------+------------+-------------+--------------------+------------+-------------------------------------------------
 hive-auto    |      10500 |       42000 | route-free-auto    | openrouter | openrouter/dots-studio/dots-3-note-preview:free
 hive-default |       5250 |       21000 | route-free-default | openrouter | openrouter/dots-studio/dots-3-note-preview:free
 hive-fast    |      10500 |       42000 | route-free-fast    | openrouter | openrouter/dots-studio/dots-3-note-preview:free
 hive-medium  |      21000 |       84000 | route-free-medium  | openrouter | openrouter/dots-studio/dots-3-note-preview:free
 hive-small   |      10500 |       42000 | route-free-small   | openrouter | openrouter/dots-studio/dots-3-note-preview:free
 hive-stt     |          0 |     4316667 | route-groq-stt     | groq       | groq/whisper-large-v3
 hive-tts     |          0 |     3080000 | route-groq-tts     | groq       | groq/canopylabs/orpheus-v1-english
```

Also confirmed against that database:

- **One enabled route per alias.** The violations query returns exactly
one row, `hive-embedding-default` with 3, which is the pre-existing
exception the integration suite already records in
`pendingMultiRouteAliases`.
- **The three batch and image flags have a healthy carrier**,
`route-free-auto`. `supports_stt` and `supports_tts` are still on the
two healthy Groq audio routes. So `/v1/batches`,
`/v1/images/generations`, `/v1/images/edits` and both voice endpoints
still find an eligible route.
- **`policy_mode` moved on no alias.** `hive-fast` is still `latency`,
`hive-small` and `hive-medium` still `pinned`, `hive-default` still
`stability`, `hive-auto` still `weighted`. Only the route name inside
`fallback_order` changed.
- **Idempotent.** Both new migrations re-applied to the same database a
second time, both clean, prices identical afterwards.
- **Integration suite green** against that database: `internal/routing`,
`internal/catalog` and `internal/litellmconfig` all `ok` with `-tags
integration`.

Running the whole `./apps/control-plane/...` integration suite at once
also produced failures in `auditworker`, `marketplace` and `tenants`.
Those are package-parallelism collisions on one shared database, not
this change: each passes on its own against the same database, and none
of them reads any of the four tables this branch touches.

## Two test-file notes a reviewer should see rather than discover

- `TestHiveFastIsPinnedToGroqAtCorrectedPrice` is **renamed** to
`TestHiveFastIsPinnedToOneRouteAtItsUnchangedPrice` and its provider and
provider_model expectations updated, because the route legitimately
moved. Its two price assertions (10500 and 42000) are unchanged and are
now the DB-level guard that this repoint did not quietly reprice three
aliases. Its comment records that these figures are no longer derivable
from the upstream's cost, which is zero, and must not be "corrected" to
match it.
- `TestSelectRouteHiveFastResolvesToGroqAtGroqPrice` in
`service_test.go` keeps a name that no longer describes the catalog. It
is a stub-driven test of the `SelectRoute` algorithm with synthetic
route fixtures; it reads neither the database nor any migration, so it
is still a valid algorithm test. Left alone rather than renamed in an
unrelated file. Push back if you would rather it were renamed.

## Residual question three, added by part two

**There is now no chat alias on a second provider.** With every
free-model alias behind one 20-requests-per-minute endpoint and no
fallbacks, a single upstream 429 takes the whole chat surface down for
that minute. Options: (a) leave it, which is what the directive
literally asks for; (b) keep one Groq chat route as a deliberate escape
hatch, cheapest candidate being the deprecated `hive-fast`, which costs
almost nothing in Groq allowance and preserves a working alias to switch
to; (c) fix #1089 first so at least the failure is a fast, correct 429 a
client library can back off from. **Recommendation: (c) then (b).** Not
done here because #1089's fix is not containable from this environment,
per the reasoning above.

@sakibsadmanshajib sakibsadmanshajib left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Adversarial review pass (the stage that had not run yet on this PR). Verdict: READY.

Findings

No blocking findings. Three non-blocking observations, none requiring changes:

  1. Low, intentional: GET /v1/audio/voices is registered without the hk_ authorizer or the tenant voice gate. Verified this is correct and low risk: Open WebUI's own voices fetch sends no Authorization header at all (vendor/open-webui/backend/open_webui/routers/audio.py, get_available_voices), so gating would silently reinstate its hardcoded alloy fallback (#996), and the handler serves a static six name roster with no database, provider, or billing contact and no request derived data. Exposure is limited to six public voice names.

  2. Info: the failure mode changes by design. With audio.tts.engine forced to openai, a gateway TTS outage now surfaces an error instead of degrading to browser speech synthesis. That silent degradation was itself the defect (#997: unmetered OS voice playback), and honest failure matches the repo's loud failure posture. No silent fallback remains.

  3. Info: environment wins for these five keys on every open-webui restart, so admin UI edits to them are reverted at next restart (compose literals are always present). This is the accepted #722 mechanism, identical to STT. One doc nit: the comment "clears AUDIO_TTS_ENGINE to keep whatever its administrator configured" cannot be done through .env alone because the compose value is a literal, it needs an override file. The same wording already ships on the preexisting STT block, so consistency rather than a defect.

Checked clean

  • Contract match against the fork: get_available_voices fetches GET {base_url}/audio/voices with no auth header and reads data["voices"] entries keyed by id/name; VoicesHandler returns exactly that shape and 405s non-GET. The speech path sends Bearer audio.tts.openai.api_key to {base_url}/audio/speech, which stays behind the existing hk_ authorizer plus tenant voice gate (untouched).
  • Alias chain: hive-tts is seeded pinned to route-groq-tts (migration 20260717_02), voice default autumn is a real Orpheus voice, LiteLLM config untouched as described.
  • No ServeMux pattern conflicts: exact pattern among exact pattern siblings (/v1/audio/speech, /transcriptions, /translations). Nothing previously served /v1/audio/voices (it 404ed into the alloy fallback), so the new route only removes a failure path.
  • Secret handling: audio.tts.openai.api_key added to SECRET_KEYS so startup logs never carry it; the paired credential refusal is tested for unset, empty and whitespace key values, persisted values untouched on refusal.
  • Tests run locally on this branch: go test ./apps/edge-api/... -count=1 -short green in the toolchain container (25 packages, audio included), and python3 scripts/test_owui_rag_env_config.py green including the new TTS pins. CI rollup green, zero unresolved threads.
  • Buglog entry carried in the PR body per protocol.

Note

Visual proof (audible read aloud playback captured on the deployed chat surface) is disclosed in the PR body as pending live capture, which can only exist once the merged deployment does. Keep that follow-up capture attached to this task after deploy per the standing visual proof directive; it is a post merge step, not a code defect.

@sakibsadmanshajib
sakibsadmanshajib merged commit 062dcd3 into main Aug 24, 2026
24 checks passed
@sakibsadmanshajib
sakibsadmanshajib deleted the fix/tts-readaloud-gateway branch August 24, 2026 13:58
sakibsadmanshajib added a commit that referenced this pull request Aug 26, 2026
…eway (#1199)

Closes #1192 (part one). Parts two and three below are decisions and
findings, deliberately not implemented here.

## What this changes

One file: `packages/openai-contract/generated/hive-openapi.yaml`,
regenerated by its own generator. No hand edits.

The generated spec had been produced from an older support matrix and
never regenerated, so it annotated 22 live endpoints
`planned_for_launch`. Since #1187 that file is served publicly and
unauthenticated at `/api/openapi.yaml` for code generators to consume,
which means every integrator who pointed tooling at it was told that
`POST /v1/chat/completions`, the gateway's primary endpoint, is merely
planned. The same applied to `/v1/embeddings`, `/v1/responses`, and all
of `/v1/files/*`, `/v1/batches/*`, `/v1/audio/*`, `/v1/images/*` and
`/v1/uploads/*`.

Regenerated with `sh
packages/openai-contract/scripts/generate-matrix.sh`, which runs
`packages/openai-contract/scripts/sync_hive_contract.py` against the
committed matrix and the pinned upstream document
(`upstream/SPEC_VERSION` records the 2026-03-28 download, so the
transform is deterministic and reproducible from the repo alone).

The diff is 27 insertions and 23 deletions: the 22 `x-hive-status`
flips, plus the `/chat/completions` `x-hive-notes` picking up the Phase
20 conditional tool-support text the matrix already carried. There is no
formatting churn, which confirms the local PyYAML matches the version
that produced the committed file. Running the generator a second time
produces byte-identical output, so it is idempotent.

## Verification

Re-running the exact comparison the console performs
(`diffSpecAgainstMatrix` in `apps/web-console/lib/api-contract.ts`):

| | before | after |
|---|---|---|
| total disagreements | 89 | 67 |
| `status_mismatch` | 22 | **0** |
| `missing_from_spec` | 67 | 67 |
| `missing_from_matrix` | 0 | 0 |

The 67 remaining are the benign ones the issue predicted and are
unchanged: 51 are `out_of_scope` and dropped from the spec on purpose by
the generator, and 16 are Hive-native endpoints that were never in the
upstream OpenAI document (part two below).

**The direction guard passes.**
`tests/unit/console-docs-contract.test.ts` is green, all 12 tests, both
with and without a populated env file. The spec operation count still
matches the `x-hive-status` annotation count exactly (97 = 97), so the
line scan that reads the spec did not stop matching after regeneration.

**The served artifact was verified, not assumed.**
`hive-web-console-prod:ci` was built from this branch and run, and `GET
/api/openapi.yaml` returned HTTP 200 with `content-type:
application/yaml`. The served body hashes to `a1aa0cda...`,
byte-identical to the regenerated file in the tree. Parsing the served
payload confirms `POST /v1/chat/completions`, `/v1/embeddings`,
`/v1/responses`, `/v1/files`, `/v1/batches`, `/v1/audio/speech`,
`/v1/images/generations` and `/v1/uploads` all now report
`supported_now`. The production image was checked specifically because
fixing only the dev image would have shipped a page that 500s in
production while every local check stayed green.

Container unit-test failures unrelated to this change:
`tests/unit/ci-web-e2e-secret-free.test.ts` fails in-container by design
(it reads a workflow file the Dockerfile never copies in). A further set
of render tests fail in the container and drop from 15 files to 8 once
an env file is supplied, so they are environment artifacts rather than
code. Only three files in the app read the contract at all
(`app/api/openapi.yaml/route.ts`, `app/console/docs/page.tsx`,
`tests/unit/console-docs-contract.test.ts`) and none of them are in the
failing set.

## Part two, a decision for the owner, not implemented here

Sixteen endpoints classified `supported_now` in the matrix have no
machine-readable schema anywhere: `/v1/rag/*` (6 operations),
`/v1/agent/tasks*` (4), `/v1/artifacts*` (3), `/v1/featuregate`, and the
Anthropic-compatible `/v1/messages` and `/v1/messages/count_tokens`. The
generator is built purely from OpenAI's upstream document, so nothing
Hive added itself can appear in its output by construction. A developer
cannot generate a client for any of them.

Three ways forward, in ascending cost:

1. **Leave it.** The spec stays a pure OpenAI-compatibility document and
Hive's own surface stays undocumented in machine-readable form. Zero
work, and the docs page keeps reporting the 16 as a visible disagreement
rather than hiding them.
2. **Author them by hand and merge them in the generator.** This is
cheaper than it looks, because the pattern already exists in this
package and is currently orphaned:
`packages/openai-contract/spec/paths/` already contains four
hand-authored Hive OpenAPI documents (`spend-alerts.yaml`,
`invoices.yaml`, `grants.yaml`, `budgets.yaml`) covering the `/api/v1/*`
control-plane surface. Nothing reads them. `sync_hive_contract.py`
consumes only `upstream/openapi.yaml` and the matrix, so those four
files are authored, committed, reviewed, and consumed by nothing. Option
2 is roughly: write 16 operations in that same style, then teach the
generator to merge `spec/paths/*.yaml` into its output. The ongoing cost
is keeping hand-written schema in step with Go handlers by review
discipline alone.
3. **Generate from the handlers.** Highest fidelity and lowest long-term
rot, but the services route with plain `net/http.ServeMux` and carry no
schema annotations, so this means adding an annotation layer or adopting
a framework. Much the largest change.

This is the difference between "an OpenAI-compatible gateway" and "a
gateway with its own documented API", so it is a product call rather
than an engineering one.

While mapping this I also found that
**`packages/openai-contract/overlays/hive-support-status.yaml` is a
third orphaned artefact**. It is a 148-action OpenAPI Overlay document
that sets `x-hive-status` per operation, it is not read by the generator
(which takes status straight from the matrix), and it is itself stale in
exactly the way the generated spec was: it still declares
`/chat/completions` `planned_for_launch`. `.wolf/buglog.jsonl` already
records it as "consumed by nothing" as of 2026-07-17. Worth deleting or
wiring up, so nobody edits it believing it has an effect.

## Part three, the matrix is itself stale against the code

Asked for because the console prints the matrix's own `generated` date.
Two problems.

**The printed date is wrong.** The matrix declares `generated:
2026-03-28`, but the file was last edited 2026-07-28 (#573), and before
that 2026-07-22 (#416) and 2026-07-17 (#352). The field is
hand-maintained and was not updated by those edits, so the console
currently prints a date roughly four months earlier than the file's real
content.

**Five route families exist in the code and are absent from the matrix
entirely**, four of which shipped after the stated generation date:

| route family | service | shipped | PR |
|---|---|---|---|
| `/v1/agent/schedules`, `/v1/agent/schedules/{id}` | edge-api |
2026-08-23 | #1081 |
| `/v1/audio/voices` | edge-api | 2026-08-24 | #1079 |
| `/v1/tenants/switch` | control-plane | 2026-05-17 | #140 |
| `/v1/admin/credit-grants`, `/v1/admin/credit-grants/{id}` |
control-plane | 2026-05-08 | #136 |
| `/v1/credit-grants/me` | control-plane | 2026-05-08 | #136 |

Per the brief I have not fixed this here, because adding matrix rows
changes runtime behaviour (see below), would need its own
integration-test additions, and would change this PR's diff again.

### The part that needs a decision quickly

The matrix is not documentation. `apps/edge-api/cmd/server/main.go:581`
wraps the entire mux in `middleware.UnsupportedEndpointMiddleware(m)`,
and `matrix.Lookup` (`apps/edge-api/internal/matrix/types.go:44`)
returns `StatusUnknown` for any method and path it cannot match exactly
or by template. The middleware answers `StatusUnknown` with a 404,
`Unknown endpoint`.

Replaying that exact lookup algorithm against the committed matrix, all
six operations of the two newest edge-api route families resolve to
`StatusUnknown`, and therefore to a 404:

```
GET    /v1/audio/voices                 -> UNKNOWN -> 404
GET    /v1/agent/schedules              -> UNKNOWN -> 404
POST   /v1/agent/schedules              -> UNKNOWN -> 404
GET    /v1/agent/schedules/{id}         -> UNKNOWN -> 404
PUT    /v1/agent/schedules/{id}         -> UNKNOWN -> 404
DELETE /v1/agent/schedules/{id}         -> UNKNOWN -> 404
(control:  POST /v1/chat/completions    -> supported_now)
(control:  POST /v1/agent/tasks         -> supported_now)
```

This is the same failure recorded in `.wolf/buglog.jsonl` as
`matrix-missing-proprietary-endpoints` on 2026-07-17, which was a demo
blocker. The guard built in response,
`apps/edge-api/internal/middleware/unsupported_integration_test.go`,
drives the real middleware against the real committed matrix, but it
enumerates only the Waves 2-4 routes. It covers neither
`/v1/agent/schedules*` nor `/v1/audio/voices`, so the same class of
regression landed again without turning anything red.

I have verified this by reading the middleware and `Lookup` and by
replaying the algorithm, not against a live stack, so please confirm
against a running edge-api before acting. If it holds, scheduled agent
tasks (#1081) and the Open WebUI voice roster (#1079) are both dead in
any deployment, and the fix is matrix rows plus adding those paths to
the integration test's table.

## Why this rotted, and the cheapest guard

Nothing regenerates the spec in CI. The only reference to this package
in `.github/workflows/ci.yml` is an unrelated lint script, so a matrix
edit has never been forced to bring the generated spec with it.

The repo already has the right pattern a few lines above, for a
different artefact:

```yaml
      - name: Codegen drift — permissions.generated.ts
        run: |
          make gen-permissions
          git diff --exit-code apps/web-console/lib/control-plane/permissions.generated.ts
```

The same shape would have caught this on the day the matrix changed:

```yaml
      - name: Codegen drift — hive-openapi.yaml
        run: |
          sh packages/openai-contract/scripts/generate-matrix.sh
          git diff --exit-code packages/openai-contract/generated/hive-openapi.yaml
```

Not added here because `.github/workflows/ci.yml` was explicitly out of
scope for this task and other lanes are live in it. Recommended as a
small follow-up.

## Two smaller findings

**The direction guard is now vacuously true, and its allowance is
stale.** `console-docs-contract.test.ts` pins the direction of status
mismatches rather than their count, which was the right call while a
known-stale direction existed. After this PR there are zero status
mismatches, so the loop body never executes. The guard is not dead (any
*new* mismatch still enters the loop and fails, including the
overstating direction it was written to catch), but the allowed
direction it whitelists is exactly the bug this PR just fixed, so that
specific regression could return silently. The test's comment also now
describes a live defect that no longer exists.

Per the brief I have not edited the test, since changing a guard in the
same PR that changes what it guards deserves a separate decision. The
one-line tightening, if wanted:

```ts
expect(disagreements.filter((entry) => entry.kind === "status_mismatch")).toEqual([]);
```

That is strictly louder than what is there now, and because the
generator stamps `x-hive-status` directly from the matrix, a mismatch is
structurally impossible unless somebody edited the matrix without
regenerating. In other words, that assertion is the drift guard, in a
place CI already runs.

**The generator writes a file the repo deleted.**
`sync_hive_contract.py` still renders `docs/support-matrix.md`, but that
file was removed from the repo by #315 when planning docs moved to the
Obsidian vault, and it is not in `.gitignore`. Every run therefore drops
a 21 KB untracked file into the tree, which a later `git add -A` would
sweep back in and silently restore a deliberately deleted document. I
removed it from my tree rather than committing it. Either drop
`render_markdown` from the generator or ignore its output, but somebody
should pick one.

## Buglog entry

To be appended to `.wolf/buglog.jsonl` on `main` in a separate
buglog-only PR after this merges, per `.claude/rules/openwolf.md`.

```json
{"id":"openapi-spec-stale-understates-gateway","priority":"P1","date":"2026-08-25","error_message":"Publicly served /api/openapi.yaml annotated 22 live endpoints planned_for_launch, including POST /v1/chat/completions, telling every integrator the gateway's primary endpoint was not shipped","root_cause":"packages/openai-contract/generated/hive-openapi.yaml is generated from matrix/support-matrix.json by scripts/sync_hive_contract.py, but nothing in CI regenerates it or fails on drift. The matrix was edited three times (#352, #416, #573) without the generated spec being regenerated, so the spec kept the statuses of a much older matrix. Invisible until #1187 started serving the file publicly and rendering a spec-vs-matrix diff on /console/docs, which surfaced 89 disagreements.","fix":"Ran packages/openai-contract/scripts/generate-matrix.sh to regenerate the spec from the committed matrix and the pinned upstream document, no hand edits. Status mismatches went 22 to 0 and total disagreements 89 to 67 (the remaining 67 are out_of_scope drops and Hive-native endpoints absent from upstream, both by design). Verified the regenerated bytes reach the served artefact by building hive-web-console-prod:ci and hashing the GET /api/openapi.yaml response body. Follow-up recommended: a CI codegen-drift step mirroring the existing permissions.generated.ts one.","tags":["contract","openapi","codegen-drift","docs","ci-gap"]}
```

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
sakibsadmanshajib added a commit that referenced this pull request Aug 26, 2026
… guard from the mux (#1203)

## Summary

Live authenticated probe against hive-demo-cf confirmed the claim in the
dispatch: `GET /v1/audio/voices` and every `/v1/agent/schedules`
operation return 404 `{"code":"unknown_endpoint"}` in production, even
with a real Supabase session, because they were never added to
`packages/openai-contract/matrix/support-matrix.json`.
`UnsupportedEndpointMiddleware` wraps the whole `/v1/` mux and answers
`StatusUnknown` with 404 before the request ever reaches the feature
gate or the handler, so both shipped features (#1079, #1081) have been
dead on the live box since they merged.

Reading the two handlers turned up two more unlisted operations in the
same family while auditing: `GET /v1/agent/tasks/{task_id}/events` and
`.../files`, real routes registered on the mux with no matrix entry at
all. All eight are added with `supported_now`.

### How I verified (before writing any code)

1. Read `UnsupportedEndpointMiddleware` and `matrix.Lookup`: confirmed
the mechanism (default case on `StatusUnknown` returns 404) and that
auth (`authSelectorMiddleware`) wraps *outside* the middleware, so an
unauthenticated probe cannot distinguish "unlisted" from "unauthorized"
— which is exactly why the earlier unauthenticated probe was
inconclusive.
2. Minted a real session via the admin one-time-token flow (magic-link
mint, no password touched) against the box's self-hosted Supabase, from
inside the `control-plane` container (it's the only container holding
`SUPABASE_SERVICE_ROLE_KEY`), for `e2e-verified@scubed.com.bd`.
3. Sent raw authenticated HTTP requests to `edge-api:8080` from inside
the docker network on the box. `GET /v1/models` (control) returned 200.
`GET /v1/audio/voices` and `GET /v1/agent/schedules` both returned:
   ```
   HTTP/1.1 404 Not Found
{"error":{"message":"Unknown endpoint: GET
/v1/audio/voices","type":"invalid_request_error","param":null,"code":"unknown_endpoint"}}
   ```
This is the literal `default:` branch body from
`UnsupportedEndpointMiddleware`, i.e. `matrix.Lookup` returned
`StatusUnknown`. Confirmed, not inferred.
4. Confirmed both endpoints are genuinely registered on the mux (`grep
mux.Handle`), so this is purely a matrix data gap, not a routing gap.

## What changed

**Data:** 8 new `supported_now` entries in `support-matrix.json`: `GET
/v1/audio/voices`, the 5 `/v1/agent/schedules` operations, and `GET
/v1/agent/tasks/{task_id}/events` / `.../files`.

**Guard, two layers (this is a recurrence — buglog entry
`matrix-missing-proprietary-endpoints`, 2026-07-17, was the same class
and the guard built then only covers a hand-typed case list nobody
extended for these two new families):**

- Kept and extended `unsupported_integration_test.go`'s hand-list test
with the 8 new cases. This is the only mechanism that can see a new
suffix inside a handler's own internal path dispatch
(`routeItem`/`routeTaskByID`-style switch) — a raw mux pattern alone
cannot.
- Added a mux-derived guard that needs no such list. `route_recorder.go`
wraps the real `*http.ServeMux`, recording every pattern registered
through it. `main()` now refuses to start (`log.Fatal`) if any `/v1/`
pattern it actually registered has zero `support-matrix.json` coverage
(`assertMatrixCoverage`), and `route_matrix_guard_test.go` exercises the
identical check in CI by calling the real registration functions
(`registerRAGRoutes`, `registerAgentTaskRoutes`,
`registerAgentScheduleRoutes`, `registerInfraRoutes`,
`registerMediaFileBatchRoutes`, the newly-extracted
`registerAudioVoicesRoute`, and `artifacts.Handler.Register`) with
lightweight fakes, mirroring the existing `gated_routes_test.go` pattern
of building a real mux from real registration code.
- Verified the new guard actually would have caught this bug: replayed
it against the pre-fix matrix via `HIVE_MATRIX_PATH_FOR_TEST` and it
fails with exactly `/v1/audio/voices, /v1/agent/schedules,
/v1/agent/schedules/`.

`registerInfraRoutes`, `registerMediaFileBatchRoutes`, the three
`gated_routes.go` functions, and `artifacts.Handler.Register` now accept
a small local interface (`httpMux` / `muxHandleFunc`) instead of the
concrete `*http.ServeMux`, so the recorder can be passed to them with
zero behavior change. Existing tests that build a plain
`http.NewServeMux()` (`gated_routes_test.go`,
`artifacts/handler_test.go`) compile unchanged — `*http.ServeMux`
already satisfies both interfaces structurally.

**Stated, known limit:** the mux-derived guard cannot see past a
registered subtree prefix into a handler's own internal method/suffix
switch — that's why the hand-list test stays rather than being deleted.
Closing that fully would mean rewriting these proprietary handlers onto
Go 1.22+ method+wildcard mux patterns (`mux.HandleFunc("GET
/v1/agent/schedules/{id}", ...)`) instead of one prefix registration
plus manual dispatch inside. Out of scope here; noted for anyone picking
this up further.

## Not fixed here (noted per dispatch instructions)

- Nothing in CI regenerates the spec/matrix from source, which is the
root cause of the drift. The repo already has the right pattern for this
in the `permissions.generated.ts` drift step; applying that pattern to
the matrix is a separate change.
- `packages/openai-contract/scripts/sync_hive_contract.py` still writes
`docs/support-matrix.md`, deleted by #315 and not gitignored, so every
run drops an untracked file that a later `git add -A` would silently
restore. Also separate.

## Deploy status

Deploys to the box are currently blocked by an unrelated failing
migration on another branch. This fix will not reach production until
that clears — merging this PR alone does not fix the live 404s.

## Test plan
- [x] `go build ./apps/edge-api/...`
- [x] `gofmt -l` clean on every changed/new file
- [x] `go vet ./apps/edge-api/...` clean
- [x] `go test ./apps/edge-api/... -count=1` — all packages green,
including the extended hand-list test and the new guard test
- [x] New guard test replayed against the pre-fix matrix via
`HIVE_MATRIX_PATH_FOR_TEST` — fails on exactly the two
originally-reported routes, confirming it would have caught this
- [x] Live authenticated probe against hive-demo-cf (see Summary) — this
is what confirmed the bug in the first place

🤖 Generated with [Claude Code](https://claude.com/claude-code)

## Buglog entry

Per `.claude/rules/openwolf.md`, carried here for the follow-up
buglog-only PR (not appended to `.wolf/buglog.jsonl` on this branch):

```json
{"date":"2026-08-25","tags":["matrix","edge-api","routing","recurrence"],"error_message":"GET /v1/audio/voices and the whole /v1/agent/schedules family returned 404 unknown_endpoint on the live box despite being fully implemented and registered on the mux","root_cause":"UnsupportedEndpointMiddleware 404s any /v1/ path with no support-matrix.json entry, checked before auth/gate/handler; the two route families shipped (#1079, #1081) without matrix entries, and the existing regression guard (unsupported_integration_test.go, added for the same defect on 2026-07-17) only covered a hand-typed case list that nobody extended for these","fix":"Added the 8 missing matrix entries (including two more found while auditing: GET /v1/agent/tasks/{task_id}/events and .../files); added a mux-derived boot-time+CI guard (route_recorder.go, assertMatrixCoverage) that fails on any /v1/ pattern the mux actually registers with zero matrix coverage, so a new route can no longer ship unlisted without a human remembering to update a list"}
```

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added support for audio voice listing and agent task event, file, and
schedule endpoints.
* Added startup validation to detect API routes missing from the support
matrix.

* **Bug Fixes**
* Improved route coverage checks to recognize exact paths, normalized
paths, and route subtrees.

* **Tests**
* Added coverage checks for registered routes, including validation that
unsupported routes are rejected.
  * Expanded regression coverage for agent and audio voice routes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
sakibsadmanshajib added a commit that referenced this pull request Sep 3, 2026
From the Go review. The exemption itself is unchanged.

The constant had been inserted between registerAudioVoicesRoute's doc
comment and the function, so godoc attached the issue #1079 paragraph to
the constant and left the function undocumented. The constant now sits
above that comment with a blank line between them.

The bigger point was two tests each naming themselves the complete
exemption set. TestOnlyTheDescriptorListIsExemptFromAuth was true when
PR #1730 wrote it and became false the moment this change exempted a
second route, and false in the direction that still passes, since its
table simply does not mention the voice roster. Rather than leave a
stale claim next to a fresh one, both move into one table,
TestAuthSelectorExemptions, which carries both exempt routes and every
negative case for both. The web tools file keeps its reachability test
and a note saying where the other half went and why it moved.

Three copies of the same middleware closure setup collapse into one
helper, exerciseAuthSelector, and the cases run under t.Run so a failure
names itself.

The real-roster test drops the Name field it decoded and never asserted,
and its comment now says plainly what it does not guard: the roster's
contents are pinned in internal/audio/handler_voices_test.go, so a
hardcoded fallback list would satisfy this test. It guards against no
roster, not against the wrong one.

Verified the consolidated table can still go red: with the voice roster
removed from the exemption, the roster case and the end-to-end test both
fail and every other case stays green. Whole package passes, gofmt and
go vet clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PGAcTcHd3PdD531LaLqXbw
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Chat has no TTS engine configured, so read aloud never reaches hive-tts, and its stored defaults point at api.openai.com

1 participant