feat(audio): add native ElevenLabs HTTP compatibility routes - #11312
diegosouzapw merged 2 commits into
Conversation
|
CI follow-up: the first run exposed the repository credential-selection inventory ratchet, not a product failure. Added the new shared helper as one classified credential-selection site and reran both focused checks locally: inventory 3/3 passed; native ElevenLabs routes 6/6 passed. The unrelated shard-4 failure is the existing static-provider catalog golden drift ( |
c38832a to
ccef3fe
Compare
|
Rebased onto current |
690f684
into
diegosouzapw:release/v3.8.50
…enerate the omni-inference skill Two gates in the Lint job, both inherited — each was hidden behind the one before it. i18n value drift: #11283 rewrote sidebar.trafficInspectorSubtitle in en.json without touching the 42 translations, so 32 locales kept serving a sentence the English no longer says. Most take the documented __MISSING__: placeholder, which makes the runtime serve the corrected English until the translation pipeline catches up. Three do not: - vi cannot take a placeholder at all — tests/unit/i18n-vi-completeness.test.ts bans any __MISSING__/__TODO__ value outright, so it needs a real translation. - pt and pt-BR are translated for real rather than placeheld, because a placeholder there means this project's own maintainer reads the sidebar in English. Each file changed by exactly one line; the JSON was not reserialised wholesale. agent-skills-sync: skills/omni-inference/SKILL.md was missing the ElevenLabs voices and speech-to-text routes added by #11312, so the generator reported one file out of date and the gate exited 2. Regenerated — purely additive, 48 lines, no deletions. Verified: check-ui-value-drift PASS, i18n:check-ui-coverage PASS (42 locales), i18n-vi-completeness 5/5, check:agent-skills-sync 46 unchanged. Refs #10692
* fix(release): restore diegosouzapw#10534 quota recovery and validate the volcengine connect bodies Two base-reds on the v3.8.50 tip, found by the release pre-flight. 1. diegosouzapw#11355 regressed diegosouzapw#10534. It replaced the per-window recovery check with an unconditional `hasActiveCooldown()` stop, which is right for an upstream-derived cooldown but also blocks the case diegosouzapw#10534 exists for: a Claude-subscription 429 persists a SYNTHETIC 1h rateLimitedUntil because the upstream sends no parseable reset. When the later poll shows every governing window has really reset with quota left, holding that synthetic cooldown just deadlocks the connection for an hour. The orphaned `windowStillExhaustedAfterRealReset()` helper and the three unused claudeExtraUsage imports that ESLint flagged were the fingerprint of this regression, not dead code: they are the two halves of the original gate. Re-wired as `isQuotaExhaustedCooldownReleasable()`, deliberately narrow — only lastErrorType "quota_exhausted" is eligible, one still-exhausted or unknown-reset window keeps the lock, and an extra-usage POLICY block stays locked even though its quota windows do look recovered in the same fetch. diegosouzapw#11277/diegosouzapw#11355 semantics are untouched (both guards still pass). Regression guard: tests/unit/provider-limits-recovery.test.ts already pinned this contract and was red on the tip. 15/15 now. 2. The three volcengine-plan connect routes read `request.json()` and handed the raw fields to a headless-browser login service after ad-hoc typeof checks (`check:route-validation:t06`, Hard Rule #7). `String(body.code ?? "")` turned 123 into "123" and an absent code into "", both reaching the service as a plausible SMS code. Now parsed with Zod schemas, before the session lookup, so a malformed body answers 400 instead of a misleading 404. New: tests/unit/volcengine-plan-connect-validation.test.ts (8 cases, red before the fix). Gate: 687 route files scanned, PASS. Also drops a genuinely dead import (formatVideoTimestamp in videoBridge.ts — only used inside the helpers module that defines it). * fix(release): drain the v3.8.50 docs/golden/GLM base-reds and restore a masked assert Second base-red batch from the release pre-flight, measured on the .113 with a clean npm ci (the devbox tree resolves eslint-plugin-react-hooks 7.1.1 from a stray pnpm store instead of the lockfile 7.0.1 and reports 925 phantom errors). Provider count 350 -> 352, one root cause behind three reds. Two providers landed this cycle (volcengine-agent-plan, volcengine-coding-plan) without regenerating the artifacts that quote the count: - docs/reference/PROVIDER_REFERENCE.md regenerated (gen:provider-reference). - README / AGENTS / llm.txt (+42 mirrors) / package.json description / 4 SVG diagrams updated, including the section heading AND the anchor that links to it, so the link does not break. - tests/snapshots/provider/translate-path.json regenerated. The diff is purely additive: 46 insertions, 0 deletions, exactly the two new providers. GLM effort tiers. diegosouzapw#11415 added the explicit glm-5.3-max tier and left two sibling vitest specs pinning the old 16-model inventory and an empty tier list for it. Aligned to the shipped contract (inventory order matches glmProvider.ts; glm-5.3-max declares ["max"]). Test-masking. Four assert reductions surfaced once the deleted-file signal was resolved. Three are legitimate and are allowlisted with their reasoning: diegosouzapw#11355 inverted the startup-cooldown contract (preserve future quota cooldowns), diegosouzapw#11280 replaced two unrolled hops with a 3-hop loop that asserts MORE, and the Gemini 3.5 Flash retirement removed the models those capability asserts described. The fourth was real masking: diegosouzapw#10960 rewrote the oneproxy status test to install a stream mock, immediately overwrite it with a passthrough to the real fetch, and assert `calls.length >= 0` — always true. Restored to assert what the test name claims (the JSON-RPC tools/call carries omniroute_oneproxy_stats and its result reaches the caller), with a scope note that it pins the MCP client contract rather than the commander wiring. Also allowlists the Gemini 3.5 Flash test deletion as _deletedWithReplacement (the model was retired by 2764812; gemini-models-parser.test.ts pins the new "excluded from the parsed list" contract), and rebaselines bundleSize 8045 -> 8461 with per-entry measurements — every entrypoint stays far below its absolute budget. * test(a2a): call the agent-card route handlers with a NextRequest PR diegosouzapw#11418 (S2 topology sanitisation) removed the hardcoded localhost:20128 from both well-known agent-card routes and made them derive the base URL from `request.nextUrl.origin` via `getBaseUrl(request)` (src/lib/wellKnown.ts). That changed the handler contract: `GET` now requires the request Next.js always passes it. Three sibling test files were never aligned and still invoked the handler as a bare `GET()`, so every case blew up with `TypeError: Cannot read properties of undefined (reading nextUrl)` before reaching a single assertion — 8 base-reds from one moved contract, not from a skill-count drift. Align the callers to the shipped contract with a local `makeCardRequest()` helper mirroring tests/unit/security-s1-s2-s4.test.ts (a Request with a defined `nextUrl`). No assertion was removed, loosened or skipped; the assert counts are unchanged and the cases now actually execute. Refs diegosouzapw#11418 * test(router-eval): give spawned CLI children an explicit DATA_DIR The router-eval CLI test spawns the CLI with spawnSync and asserts stderr stays empty. NODE_TEST_CONTEXT is inherited by those children, so since diegosouzapw#10432 (guard diegosouzapw#10428) resolveWritableDataDir() detects a test context with no DATA_DIR and warns on stderr before falling back to a throwaway dir - 194 chars that broke three cases. Pass an isolated DATA_DIR in the child env (the resolution the guard message itself prescribes) instead of loosening the assertions. * test(quota): freeze the clock in the GLM absolute-ISO-reset regression test The diegosouzapw#11353 regression test pins an ABSOLUTE upstream reset instant (2026-08-29 21:01:21) in its production fixture body but measured the resulting cooldown against the real wall clock. The remaining window therefore shrank every day: from 2026-08-25 it dropped under the 5-day floor the two assertions use, and past 2026-08-29 it would parse to null and collapse onto the 24h WEEKLY_QUOTA_COOLDOWN_MS default - a guaranteed future red. The shipped parser (parseIsoDateTimeResetMs / parseDayGranularityResetMs / buildWeeklyQuotaFallback / checkFallbackError) is correct: it returned the real multi-day reset, just measured from today instead of the fixtures NOW. Freeze Date at NOW via node:test mock timers in the two time-dependent cases so they assert the parser rather than the calendar. No assertion weakened, no production code touched. * test(providers): use a non-reserved prefix in the CC-compatible node create cases 93da24c ("fix(providers): reject reserved provider prefixes on compatible-node create/update") made createProviderNodeSchema reject any prefix that is a built-in REGISTRY id or alias. "cc" is the alias of the built-in `claude` provider, so the two provider-nodes create cases in cc-compatible-provider.test.ts started getting a 400 schema rejection before the route ever reached its feature-flag gate (403) or the create path (201) — the guard PR updated its own tests but missed this sibling file, leaving a base-red on release/v3.8.50. The operator-chosen prefix is incidental to what these cases assert (the ENABLE_CC_COMPATIBLE_PROVIDER gate, the dedicated anthropic-compatible-cc- id prefix, baseUrl sanitization and the nulled modelsPath), so switch it to a non-reserved "cc-proxy". No assertion was removed or loosened. * fix(quality): drain the three v3.8.50 inventory/coverage base-reds All three guards were drifting behind legitimate cycle growth, not catching a defect. Nothing was weakened: no assertion removed, no floor lowered, no blanket-allow added. providers-constants-split: APIKEY_PROVIDERS 231 -> 233. The delta is exactly the two Volcano Ark plan providers (volcengine-agent-plan, volcengine-coding-plan) added to the regional family in d732cf6. The invariant the guard exists for still holds, measured on the tip: 233 merged keys, 233 unique, family sum 233 (gateways 92 + frontier-labs 25 + inference-hosts 29 + enterprise-cloud 17 + regional 43 + specialty-media 27) with an empty cross-family duplicate set and an empty symmetric difference between the merged object and the family union - so the six files are still a strict partition, no loss and no dup. openapi-coverage: the operation floor (34.6%) is untouched. The cycle grew the denominator 985 -> 1002 while covered only moved 343 -> 345 (34.4%). Fixed by DOCUMENTING five real public operations rather than moving the floor, taking it to 350/1002 = 34.9%: GET /api/health, GET /api/v1/voices, POST /api/v1/speech-to-text, POST /api/v1/text-to-speech/{voiceId} and GET /api/v1/explain/routing. Each entry was written from the route source (auth mode, path-param pattern, limit clamp, upstream relay behaviour and the 400 / 401 / 429 branches), not from memory. hard-session-lease-bypass-inventory: three new connection-query sites classified, none silenced. open-sse/services/combo.ts (readConnectionForCooldownGate) reads the row backing the pre-dispatch persisted-cooldown gate, so it sits on the routing path and joins the class-B list next to combo/providerWildcard.ts and autoComboCandidates.ts. src/lib/providers/volcenginePlanBinding.ts and src/lib/providers/volcPlanAutoSyncBackfill.ts are connection persistence, not dispatch - the first resolves update-vs-create during connect, the second is a one-shot boot backfill of a providerSpecificData flag with no upstream call - so both stay class C alongside oauth/connectionPersistence.ts. * fix(release): drain two v3.8.50 base-reds (volcengine vision metadata, antigravity BYOP contract) Two unrelated real reds on the release tip: * fix(providers): flag MiniMax M3 as multimodal on the Volcengine Ark plans. d732cf6 ("feat(volcengine): add Ark plan providers") added volcengine-agent-plan/minimax-m3 and volcengine-coding-plan/minimax-m3 without supportsVision, breaking the LEDGER-4 invariant that every minimax-m3 registry entry except PromptQL (text-only upstream) is flagged multimodal. Every other provider carrying the model (opencode-zen, opencode-go, bazaarlink, ollama-cloud, codebuddy-cn, trae) sets it. Registry metadata defect, not a stale test. * test(antigravity): align the empty-projectId onboarding test to the contract shipped by diegosouzapw#11284/diegosouzapw#11358 (6de542b). That change made an onboardUser 200 whose body carries NO cloudaicompanionProject mean Google BYOP — no project was created and none ever will be — so it short-circuits before the retry loadCodeAssist. The older test still mocked onboardUser with the bare { done: true } BYOP shape while asserting the retry path, so it pinned a contract that was deliberately moved. The mock now returns a real onboarding-success body; every assertion is kept, and the id in the onboard body deliberately differs from the expected one so the test still proves the projectId came from the retry discovery. Refs diegosouzapw#11284 * fix(cli): do not treat a tmpfs mount as proof a config path reaches the host hasBindMountAt() accepted ANY mount as evidence that a would-be CLI config write reaches the operator's host: it matched on the mount point alone and never looked at the filesystem type. An in-memory mount therefore cleared the ephemeral flag, so guardCliConfigWrite() let the write through and both POST /api/cli-tools/apply and the dashboard's guide-settings writer answered 200 instead of the safe 422 that diegosouzapw#10057 added. That is the exact case the guard exists to refuse, and the worst one: a container running with `--tmpfs /tmp` (or a home on tmpfs) loses the file even before the container is recreated, while the UI reports success. Parse the filesystem type from mountinfo (the field after the lone "-" separator) and skip mounts backed by RAM or kernel state. Real bind mounts (ext4/xfs/nfs/virtiofs/fuse.*) still count, including one nested under a tmpfs path, so the compose `host` profile is unaffected. A line carrying no separator proves nothing and is skipped too. Regression cover added to tests/unit/container-env-detect.test.ts; this also un-reds tests/unit/cli-tools-apply-container-422.test.ts and tests/unit/api/cli-tools/apply-container-guard.test.ts, which were failing on any box whose /tmp is a tmpfs. * chore(docs): commit the next-dev agent-rules block into AGENTS.md `next dev` writes and re-adds this block (see node_modules/next/dist/server/lib/generate-agent-files.js), so leaving it out of a diff only recreates the uncommitted change on the next dev run. Committing it keeps the working tree clean, which is what the block's own note prescribes. * docs(changelog): aggregate the 363 changelog.d fragments into [3.8.50] Release reconciliation (Phase 0a.1). `scripts/release/aggregate-changelog.mjs` folds each changelog.d/<section>/*.md fragment into its heading in the living [3.8.50] section and deletes the fragment, which is the whole point of the fragment convention: two PRs never touch the same file, so the CHANGELOG never conflicts mid-cycle and no bullet is eaten by a merge auto-resolve. Section bullets 731 -> 1041. The remaining uncovered commits (mostly merges from diegosouzapw#11397 onward, which landed without a fragment) are reconciled separately. * docs(changelog): reconcile the v3.8.50 section with the cycle's uncovered commits Adds 131 consolidated bullets (45 features, 71 fixes, 15 maintenance) covering the ~490 user-facing commits and the ~100 chore/ci/test/refactor/docs commits that landed in the cycle without a CHANGELOG entry, grouped by subsystem and citing their PR references. Uncovered report: 594 -> 175 (the remainder are commits carrying no #N in their subject, which the matcher can never resolve; they are covered in prose). * docs(changelog): date the 3.8.50 header, inject contributors and sync the 42 i18n mirrors * fix(ci): raise the unit-shard heap ceiling to 8192 MB to stop the SIGABRT OOM The 8 unit shards run under V8 coverage instrumentation, which retains far more memory than the bare suite. With the 4096 MB ceiling they began aborting with exit 134 ("Ineffective mark-compacts near heap limit") at ~4086 MB as the provider catalogue grew during the v3.8.50 cycle: every test in the shard passed and the process died at the end, which reads as a test failure without being one. Aligns test:unit:ci:shard and test:unit:serial with the 8192 MB the non-sharded variants already use. GitHub-hosted runners have 16 GB, so the headroom is real. Validated by the CI run on this commit — the shards are the gate. Refs diegosouzapw#10692 * fix(ci): set NODE_OPTIONS on the unit-shard step and prune stale eslint suppressions The previous heap bump only touched test:unit:ci:shard, i.e. the node the shard script spawns. The process that actually runs out of memory is the `c8` wrapper around it — it aggregates ~577 MB of raw V8 coverage JSON — so the ceiling stayed at the V8 default (~4 GB) and the shards kept aborting at ~4083 MB, byte for byte the same failure. Setting NODE_OPTIONS on the step covers c8 and every child, which is the pattern the coverage-merge job already uses. Also prunes three eslint suppression entries whose violations no longer exist: videoBridgeContactSheet.ts and videoBridgeRuntime.ts (no-unused-vars, fixed during this cycle) and cli-oneproxy-commands.test.ts (no-explicit-any 14 -> 13, a consequence of restoring the real mock in that test). Stale entries make `npm run lint` exit 2 with 'There are suppressions left that do not occur anymore'. Pruned and verified on an uncontaminated checkout, not the devbox. Refs diegosouzapw#10692 * fix(ci): drain four inherited base-reds in packaging, electron and integration tests All four predate this session — each reproduces identically on f95b03d (2026-08-24), so none is a cycle regression. Draining them here because the release pre-flight is where inherited reds get resolved. 1. Package Artifact: the job runs `build:cli`, which assembles dist/ but never writes dist/BUILD_SHA — only `build:release` does, via write-build-sha.mjs. The diegosouzapw#10427 provenance guard inside check:pack-artifact then rejects the artifact as untraceable, and rejects it even under OMNIROUTE_ALLOW_CANARY_BUILD. The job's build+validate pair was structurally incompatible and failed 100% of the time. Stamps the SHA between the two steps. 2. Electron Package Smoke: electron/package.json's build.files allowlist enumerates each lib/*.js by hand and never got lib/loginHeaderCapture.js, added alongside its require() in diegosouzapw#9984. The file therefore stayed out of app.asar and the packaged app died at startup on 'Cannot find module ./lib/loginHeaderCapture'. 3. proxy-pipeline: the breaker assertion grepped chat.ts for executeChatWithBreaker(, but that call moved behind the chatDispatch.ts seam. Rather than drop the check, it now pins both hops — chat.ts dispatches through the seam and the seam calls the breaker — so the extraction cannot silently take the breaker off the path. 4. skills-pipeline: diegosouzapw#9058 began encoding skill tool names as omr_skill_<base64url> because providers require ^[a-zA-Z0-9_-]+$, and these assertions still expected the raw name@version. They now derive the expected name from encodeSkillToolName(), the same helper production uses, so the test tracks the contract instead of duplicating it. Only the assertions about names on the wire were converted; the identifiers passed straight to skillExecutor.execute() stay raw, because those are not encoded. Integration suite for these two files: 54/55. The one still red — 'web_search fallback preserves Responses API output' — is a separate pre-existing defect, deliberately left failing rather than papered over: on the /v1/responses path resolveSearchCredentials() returns null for the seeded serper-search connection, so executeWebSearch.ts:185-200 falls through to the cheapest fallbackOnly provider (duckduckgo-free) and the results come back empty. The sibling chat-path test seeds identically and does resolve serper-search. Needs its own investigation. Refs diegosouzapw#10692 * test(vitest): realign five sibling-test contracts unmasked by the green mcp shard None of these are cycle regressions. The Vitest job runs test:vitest (mcp shard) then test:vitest:ui; the mcp shard was failing on a missing glm-5.3-max and aborted the job before the ui shard ever ran. Fixing that shard this cycle unmasked 34 ui failures that had been broken since 18-23 Aug — four separate PRs that moved a contract and updated their own tests but not their siblings. - ProviderCard gained useRouter() in diegosouzapw#10448; four test files render it without mocking next/navigation and died on 'invariant expected app router to be mounted'. The sibling created alongside diegosouzapw#10448 already had the mock — it just was not applied to the other four consumers. - SkillCoverage gained a required config category. The four fixtures in agent-skills-page still described only api/cli, so the component read config.have off undefined. Values were chosen per scenario rather than pasted: full coverage gets 2/2 so its bar stays emerald, the amber fixture gets 3/4 so it stays amber. CoverageBar renders api -> config -> cli, so the new bar lands in the MIDDLE and the cli assertions moved from index [1] to [2]; without that the cli checks would have passed while measuring the config bar. The aria test now pins all three bars. - CliAgentsPage hardcoded AGENT_IDS, which had already drifted once (6 -> 8 with omp/letta, per its own comment) and drifted again with prime-agent (diegosouzapw#11166). It is now derived from CLI_TOOLS. This is why an agent missing from that list is not cosmetic: it never enters the status map, defaults to not_installed, and adds a phantom card to the filter and count tests. Deriving keeps the fixture in sync by construction instead of waiting for the next agent. - claudeTlsClient asserted proxyUrl was undefined inside a test literally named 'falls back to env var when per-call proxyUrl not provided' — it pinned the old behaviour where testOverride bypassed proxy resolution. diegosouzapw#10910 moved resolution ahead of the override on purpose ('so test overrides and the real path both see it'), so the assertion now checks the fallback the test name promises. test:vitest:ui goes from 34 failures to 14. The remaining 14 sit in six files none of this commit touches (AutoComboCatalog, CoolingConnectionsPanel, ProxyRegistryManager x2, connectionsSearchFilter) plus one claudeTlsClient case that passes in isolation and only fails in the full run — i.e. cross-file pollution. They need a clean environment to judge: this devbox resolves part of its tree through a stray pnpm store and has already produced one phantom failure count this cycle. Refs diegosouzapw#10692 * fix(i18n): unstale trafficInspectorSubtitle across 32 locales and regenerate the omni-inference skill Two gates in the Lint job, both inherited — each was hidden behind the one before it. i18n value drift: diegosouzapw#11283 rewrote sidebar.trafficInspectorSubtitle in en.json without touching the 42 translations, so 32 locales kept serving a sentence the English no longer says. Most take the documented __MISSING__: placeholder, which makes the runtime serve the corrected English until the translation pipeline catches up. Three do not: - vi cannot take a placeholder at all — tests/unit/i18n-vi-completeness.test.ts bans any __MISSING__/__TODO__ value outright, so it needs a real translation. - pt and pt-BR are translated for real rather than placeheld, because a placeholder there means this project's own maintainer reads the sidebar in English. Each file changed by exactly one line; the JSON was not reserialised wholesale. agent-skills-sync: skills/omni-inference/SKILL.md was missing the ElevenLabs voices and speech-to-text routes added by diegosouzapw#11312, so the generator reported one file out of date and the gate exited 2. Regenerated — purely additive, 48 lines, no deletions. Verified: check-ui-value-drift PASS, i18n:check-ui-coverage PASS (42 locales), i18n-vi-completeness 5/5, check:agent-skills-sync 46 unchanged. Refs diegosouzapw#10692 * test(vitest): pay heavy module imports at collection time, not out of the per-test budget The remaining ui-shard reds were one class, not six bugs: every one of them did `await import(<heavy component>)` INSIDE an `it()`, so Vite's transform of the dependency graph was billed to that test's timeout. Measured costs against the budgets they had to fit in: ProxyRegistryManager 86s import vs 30s / 60s / 5s budgets (render itself: 567ms) claudeTlsClient ~12s import vs 5s default useProviderConnections 1050-line hook, whole dashboard graph, vs 5s default That is why they looked like cross-file pollution: on an idle box the import squeaked under the limit, and under the ui suite's 20 parallel workers it did not. Running claudeTlsClient ALONE on a loaded box reproduces it — the trigger is CPU contention, not a neighbouring file. The sibling chatgptTlsClient/grokTlsClient tests import the same graph and never fail, because they import statically at module scope, where the cost falls on the collection phase which has no per-test budget. Every fix here does the same: static import or a beforeAll with its own budget. AutoComboCatalog also explains its own blast radius: the timeout aborted inside an open act(), leaking an unbalanced act scope that then failed the file's three remaining tests in ~20ms with 'overlapping act() calls'. One slow import, four reds. CoolingConnectionsPanel is the one production change. It imported providerText from the ../providerPageHelpers barrel, but that symbol is DEFINED in the ../providerCredentialText leaf and only re-exported by the barrel — which drags providerRegistry (352 providers) and the rest of the provider-page graph into a "use client" component for one string helper. Verified before accepting: the component used nothing else from the barrel, the barrel has no top-level side-effect to lose (the empty-registry hazard this repo has hit before does not apply), typecheck:core is clean, and the panel's first test drops from ~4s to 95ms. The import was suboptimal, never broken — the screen was not failing for users. No assertion was weakened anywhere. expect() counts are unchanged (25/25, 4/4) or up by one (AutoComboCatalog 11 -> 12); the diegosouzapw#8855 autofill sentinels, the data-1p-ignore / data-lpignore guards and the dead-status round-trip are intact. The diegosouzapw#5918 TDZ guard was proven still live by mutation, not by absence of red: moving useProxyBatchOperations(load) above its const reproduced 'ReferenceError: Cannot access load before initialization' in 207ms, then the production file was restored (diff empty). tests/unit/ui under load: 17 failed files / 45 failed tests -> 4 failed files / 4 failed tests, none of them these. The four left are compression-guidance-7530, compressionPanel, compressionUltraTier and lobe-provider-icons-stepfun, untouched and uninvestigated. Refs diegosouzapw#10692 * test(integration): realign three suites to security and version contracts that moved All three are the sibling-test gap again: a PR moved a contract, updated its own tests, and left these behind. None is a production defect — in two of the three the production side is a deliberate security fix. v1-contracts-behavior (4 failures, one cause): the job env sets INITIAL_PASSWORD, which makes isAuthRequired() true, and diegosouzapw#9320 (b07182c) made the /v1 catalogue gate on-by-default instead of opt-in via settings.requireAuthForModels. The four contract reads were calling the catalogue routes with no credential and correctly getting 401. Bisected the job's four env vars to confirm INITIAL_PASSWORD alone reproduces it (5 pass / 4 fail with it, 9 / 0 without). The tests now send a Bearer token; the shape assertions are untouched, and the auth contract itself stays owned by tests/unit/v1-models-auth-leak-9320.test.ts rather than being duplicated here. opencode-config-startup: two independent drifts. OPENCODE_VERSION was pinned to 1.18.8 while the installed opencode-ai is 1.18.18 (Dependabot 7f69589, diegosouzapw#10626) — now read from require("opencode-ai/package.json").version, which is exactly as strict but cannot drift on the next bump. And the no-limit-metadata case asserted limit === undefined, but diegosouzapw#11054 made the generator always emit a limit; it now pins the actual fallback {context: 128_000, output: 8_192} instead of an absence. memory-pipeline: diegosouzapw#11040 (GHSA-cpv3-xr7r-xf8q) made the resolved caller principal always win over a caller-supplied apiKeyId, so a spoofed id can no longer write into another principal's store. That PR updated the unit sibling but not this one. The test now asserts the stronger property — and deliberately not just the absence: the spoofed principal's store is empty AND the caller can still read the entry, which proves the write was redirected rather than dropped and keeps the emptiness check from passing vacuously with a disabled store. (The old assertion was count === 0, which a switched-off memory store would satisfy.) Assertion counts: 43 -> 43, 13 -> 14, 76 -> 81. Nothing weakened or removed. Verified: 24/24 pass, with and without the CI env vars. Refs diegosouzapw#10692 * fix(ci): drain the electron packaging regression, the models-catalog e2e assertion and 10 integration reds Electron Package Smoke — a packaging defect that had been hidden behind another packaging defect for nine days. Once the loginHeaderCapture fix let the main process start, the server underneath died on 'Cannot find module next': resources/app/server.js shipped without resources/app/node_modules. electron-builder discards the ROOT node_modules in code, not by configuration — app-builder-lib/out/util/filter.js:42 has a hard-coded `if (relative === "node_modules") return false` that runs before any filter pattern. The second extraResources entry pointing INTO ../.build/electron-standalone/node_modules is what sidesteps it, because those relative paths are never equal to "node_modules". diegosouzapw#10325 removed that entry as an apparent duplicate and flipped the test to assert "exactly once", freezing the regression as if it were the contract. Restored, and the unit guard now pins both entries — proven by mutation: reverting package.json to the post-diegosouzapw#10325 shape fails the guard 3/4, restoring it passes 4/4. group-b-quota-plans-config — the assertion was impossible to satisfy on ANY route, and the page was never broken. layout.tsx hands the whole message catalogue to NextIntlClientProvider, React serialises that prop into the RSC payload, and en.json carries "Internal Server Error" twice, so page.content() always contains it: probing /dashboard, /dashboard/costs, /dashboard/settings and /login showed the string present with every page rendering fine, and a pageerror probe on the failing run captured zero client exceptions. This is the same trap that killed the sibling not.toContain("500") in fc77100 ("raw HTML is unreliable") — that one was removed, this one was kept. Now asserts on rendered text, which still catches a real error boundary. The pageerror capture stays: the CI failure carried no stack trace, which is why it was misread twice. Integration — 10 of the 14 shard-2 reds, all sibling-test gaps behind security fixes: monitoring health now takes a Request and requires management auth (GHSA-mvf8-qc78-5mxm); the OAuth import routes moved to requireManagementAuth (GHSA-mg76) — the test accepts both guard shapes and gained a stronger anchor that every exported handler awaits a guard on its own request, mutation-verified; skill tool names are derived from encodeSkillToolName() and the fake upstream now returns the encoded name so decodeSkillToolName() is exercised too; previous_response_id now fails closed (diegosouzapw#10262); proxy_logs persist as an async batch (diegosouzapw#11182) so the test flushes first; providerQuotaOverrides joined GET /api/resilience (diegosouzapw#9871); the reasoning fixture used a model that stopped being thinking-incompatible, replaced and pinned with a premise assert so it cannot rot silently again. A vacuous assert.ok(true, "all 10 streams completed without hanging") was replaced with real anchors — content must arrive on every stream and the active Timeout count must not grow. Four are deliberately left red rather than aligned, each now tracked: diegosouzapw#11551 (the /v1/models after() wiring is dead — the route passes a third argument to a two-parameter function and catalogCache never imports after, so the diegosouzapw#8728 contract is unimplemented), diegosouzapw#11552 (~27% of requests emit an extra discarded upstream call; the delivered distribution is exactly 0.70, so weighted routing is correct and the waste is the real finding), the fixed-account combo pin (aligning it would destroy the per-step attribution the test exists for), and the web_search fallback already tracked as diegosouzapw#11524. Package Artifact — the provenance stamp I added last round used git rev-parse HEAD, which under pull_request is the ephemeral merge commit and therefore never an ancestor of the release branch. Now takes the PR head sha. Refs diegosouzapw#10692 * test(ui): unmount before asserting so the auto-sync timer cannot outlive the test Vitest went red on 'synchronizes upstream models only when autoFetchModels is explicitly true' with new URL throwing inside a fetch dispatched from Timeout._onTimeout (useProviderModels.ts:69). It is intermittent: green in the two previous CI runs, green every time in isolation, red only under the ui suite's 20 parallel workers. The hook schedules its auto-sync in a setTimeout whose callback only checks the flag on entry — and that flag stays false while the component is mounted. Both tests asserted first and unmounted last, so under contention the timer escaped the test window, fired after afterEach had already run vi.unstubAllGlobals(), and reached the REAL fetch with a relative URL. Unmounting before the assertions closes the window: cleanup flips , the callback returns early, and the calls already recorded on fetchMock are still there to assert against. No assertion changed. Not a regression from this cycle — the file's last change is fd76271 (diegosouzapw#10603). Fixed rather than tracked because an intermittent red in a blocking job is worse than a permanent one: it teaches people to re-run instead of to look. Refs diegosouzapw#10692 * fix(search): prefer the configured search connection over duckduckgo-free When no explicit provider is requested and the auto-selected cheapest provider has no credentials, executeWebSearch ran the fallbackOnly loop first. duckduckgo-free (costPerQuery 0, authType none) always won there with an empty credentials object, so a configured paid connection such as serper-search was silently ignored and the caller got success:true with zero results. Move the sweep for other credentialed regular providers ahead of the fallbackOnly loop (and exclude fallbackOnly ids from it, so a free last-resort provider never outranks a configured one on cost). The fallbackOnly loop stays as the true last resort. The chat path only appeared correct because duckduckgo-free happened to fail there and handleSearch retried the alternate provider; on /v1/responses it "succeeded" with no results. Closes diegosouzapw#11524 * fix(api): restore the after() injection point for the /v1/models SWR refresh `/v1/models` has passed a third argument to `getUnifiedModelsResponse()` (`{ scheduleBackgroundRefresh: (task) => after(task) }`) ever since diegosouzapw#10198, but diegosouzapw#9199 had already removed the parameter: the function takes two, so the object was silently dropped and `catalogCache` kept scheduling the stale-while- revalidate rebuild with `setTimeout(..., 0)`. The builder is overwhelmingly synchronous under the single-threaded App Router, so it pinned the event loop before the stale response was flushed — the diegosouzapw#8728 guarantee did not exist. - `catalogCache` now imports `after` from `next/server` and exposes `defaultBackgroundRefreshScheduler`, which defers to `after()` and falls back to a macrotask outside a Next request scope (instrumentation warm-up, tests). - `resolveCachedCatalogResponse` takes `scheduleBackgroundRefresh` and `getStaleWhileRevalidateMs` on its existing options object. - `getUnifiedModelsResponse` accepts the options object the route already passes and propagates the scheduler. The excess argument was invisible to CI: `tsconfig.typecheck-core.json` is a curated 27-file allowlist, `check:dashboard-typecheck` only covers `src/app/(dashboard)`, and `next.config.mjs` sets `ignoreBuildErrors: true`. Closes diegosouzapw#11551 * fix(sse): stop re-summarizing a universal handoff that never parses A universal handoff whose summary comes back unparseable persists nothing, so shouldGenerateUniversalHandoff keeps answering "generate" and the very next model switch in the same session re-issues the same full-history summarization call and discards the answer again — forever. With a switch-heavy combo strategy (weighted, random, round-robin, p2c) the models alternate on almost every turn, so that background call lands on a large fraction of requests: an upstream call whose response nobody reads is real money on a paid provider and real quota on a metered one. Measured on the weighted 70/30 matrix, that inflated the observed openai share to 0.895 where routing actually delivered 0.70, and at the unit level 199 of 200 model switches issued a fresh discarded summarization call. Back off per (session, combo) after an answer that is not a usable handoff: exponential 5min -> 1h, cleared on the first successful generation, capped at 500 tracked keys. Deliberately narrow — a transient upstream failure (!response.ok) is NOT tracked, so it still retries on the next switch, which is the behavior the context-relay path already depends on. After the fix, same harness at n=200: 201 upstream calls for 200 requests (1 extra, 0.5%), and the measured openai share equals the delivered share (0.725). The unit guard drops 199 discarded calls to 1. Refs diegosouzapw#11552 * fix(sse): honor per-step connection pins instead of rotating accounts on fallback A combo step pinned with an explicit `connectionId` (or a request pinned via `x-omniroute-connection`) is an operator instruction, not a hint. The generic account-fallback branch in handleSingleModelChat excluded the pinned connection after an upstream failure and re-selected a sibling account of the same provider, so a priority combo repeating one provider/model with two different fixed accounts ran both attempts under the FIRST step: the second step, with its own pin, never executed and per-step attribution (comboStepId / comboExecutionKey) was wrong. Gate the rotation on `!hasForcedConnection`, matching the antigravity stream-readiness, pre-response-timeout and account-semaphore branches that already let pinned steps fall through to combo orchestration. Cooldown recording via markAccountUnavailable is unchanged, and unpinned selection still excludes burned connections. Refs tests/integration/combo-routing-e2e.test.ts * fix(ci): point the pack-artifact provenance gate at the branch under test The guard added for diegosouzapw#10427 checks ancestry against origin/main by default. That is the right ref at publish time (npm-publish.yml runs on main), but in a pull_request context it can never hold: while the PR is open its head is by construction not an ancestor of main, and the shallow checkout does not even bring origin/main into the local graph, so the probe answers false regardless. The job therefore failed 100% of the time and only became visible now that it stopped being cancelled behind Build. Pre-merge the one checkable invariant is that the stamp matches the branch under test, so resolve the ref from refs/pull/<N>/head — which exists on origin even for fork PRs, unlike head.ref, which only exists on the author's repository. * test(dashboard): open the proxy toolbar overflow menu before bulk assign PR diegosouzapw#9870 moved the proxy-registry bulk-assign action into the toolbar's "More actions" (…) overflow menu, which renders its items only while open. The e2e smoke flow still clicked the testid directly, so the locator never resolved and the test burned its full 180s budget. Open the menu first and assert it is visible before clicking; every existing assertion is unchanged. * fix(dashboard): restore expert-mode manual model entry in the combo builder PR diegosouzapw#8285 (global model search) replaced the expert-mode "Manual model" block positionally with the new GlobalModelSearchPanel, dropping the only way to type a provider/model pair by hand in expert mode. The supporting state and handlers (manualModelInput, manualModelError, manualModelHasDuplicate, handleAddManualModel) survived as dead code, so neither typecheck nor lint flagged the loss. Re-render the block above the search panel, unchanged from its pre-diegosouzapw#8285 form. Regression guard: tests/e2e/combos-flow.spec.ts "expert mode shows a single-page combo form with manual model entry", which had been failing with a 180s timeout on locator.fill for combo-manual-model-input. * test(api): accept the post-diegosouzapw#9320 auth gate on the /v1/models e2e check b07182c (diegosouzapw#9320) inverted the catalog auth rule: /v1/models now requires auth whenever management auth is configured, unless requireAuthForModels is explicitly false. The e2e harness boots with INITIAL_PASSWORD set, so the endpoint has been answering 401 since 2026-08-04 and this check has been red ever since — invisible only because the job kept being cancelled behind Build. Mirror the sibling /api/providers check in this same file: assert the catalog shape whenever the catalog is actually served, and otherwise pin the auth gate by status AND error type, so a 401 from an unrelated misroute cannot pass for the deliberate one. --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: Xiangzhe <diegosouza.pw@gmail.com> Co-authored-by: Xiangzhe <bakryun0718@proton.me> Co-authored-by: TheDemonTuan <nguyenviettuanbp@gmail.com>
…uzapw#11312) Validated on a 17-PR combined board: elevenlabs-native-routes + hard-session-lease-bypass-inventory (9/9) within the board's 287/287, typecheck:core clean. Native ElevenLabs compatibility routes (voices, TTS, STT) reusing the stored credential via quota-preflight, sent only as xi-api-key; client authorization headers never forwarded. Closes diegosouzapw#10556. Thank you @RaviTharuma!
…enerate the omni-inference skill Two gates in the Lint job, both inherited — each was hidden behind the one before it. i18n value drift: diegosouzapw#11283 rewrote sidebar.trafficInspectorSubtitle in en.json without touching the 42 translations, so 32 locales kept serving a sentence the English no longer says. Most take the documented __MISSING__: placeholder, which makes the runtime serve the corrected English until the translation pipeline catches up. Three do not: - vi cannot take a placeholder at all — tests/unit/i18n-vi-completeness.test.ts bans any __MISSING__/__TODO__ value outright, so it needs a real translation. - pt and pt-BR are translated for real rather than placeheld, because a placeholder there means this project's own maintainer reads the sidebar in English. Each file changed by exactly one line; the JSON was not reserialised wholesale. agent-skills-sync: skills/omni-inference/SKILL.md was missing the ElevenLabs voices and speech-to-text routes added by diegosouzapw#11312, so the generator reported one file out of date and the gate exited 2. Regenerated — purely additive, 48 lines, no deletions. Verified: check-ui-value-drift PASS, i18n:check-ui-coverage PASS (42 locales), i18n-vi-completeness 5/5, check:agent-skills-sync 46 unchanged. Refs diegosouzapw#10692
Summary
GET /v1/voices,POST /v1/text-to-speech/{voiceId}, andPOST /v1/speech-to-text.elevenlabscredential through OmniRoute's quota-preflight path and sends it only asxi-api-key.voiceIdto one safe URL segment before constructing the fixed-host upstream URL./v1/*CLIENT_API authz and CORS policy authoritative; client authorization headers are never forwarded.Related Issues
Validation
npm run lint— focused ESLint passed for every changed TypeScript file; full gate runs in CIrelease/v3.8.50atac02c5b42; focused checks rerun afterwardnpm run check:changelog-integritygit diff --checkPost-rebase focused commands:
node --import tsx/esm --test tests/unit/elevenlabs-native-routes.test.ts tests/unit/hard-session-lease-bypass-inventory.test.ts— 9 passednpm run check:changelog-integrity— passedThe current release base has a pre-existing
typecheck:core/typecheck:noimplicit:corefailure inopen-sse/translator/response/openai-responses/pureHelpers.ts(implicit-any parameters). This PR does not touch that file.Tests Added Or Updated
tests/unit/elevenlabs-native-routes.test.tstests/unit/hard-session-lease-bypass-inventory.test.tsRegression coverage verifies URL/query forwarding, stored
xi-api-keyuse without leaking it in responses, JSON TTS requests, multipart STT uploads, binary audio responses, upstream error/status preservation, missing credentials, traversal rejection, and classification of the shared credential-selection site.Coverage Notes
The new route-level test exercises all three handlers plus the shared proxy's successful JSON/binary/multipart paths, missing-credential branch, upstream non-2xx pass-through, and unsafe-voice rejection. CI owns the full coverage ratchet and production build.
Reviewer Notes
No migrations, dependencies, settings, or feature flags. The proxy host and native paths are fixed constants. Response headers are allowlisted rather than copied wholesale. Transport exceptions pass through OmniRoute's sanitized error helper.