Skip to content

feat: headroom integration + resource optimizations for container stability - #6572

Closed
oyi77 wants to merge 166 commits into
diegosouzapw:mainfrom
oyi77:feat/headroom-resource-optimizations
Closed

oyi77 wants to merge 166 commits into
diegosouzapw:mainfrom
oyi77:feat/headroom-resource-optimizations

Conversation

@oyi77

@oyi77 oyi77 commented Jul 7, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds Headroom (token-saver proxy) integration and critical resource optimizations to fix container instability under multi-agent load.

Changes

1. Headroom CLI baked into container ()

  • Installs Python 3, pipx, and globally
  • Enables and endpoints to work out of the box

2. Resource limits & V8 tuning ()

Setting Before After
Memory limit Unlimited 8 GB
CPU limit Unlimited 4 cores
V8 heap (max-old-space-size) 1 GB 6 GB
Healthcheck timeout 5s 10s
Healthcheck start_period 15s 30s

3. Local font build fix ()

  • Switches from to with bundled Inter WOFF2
  • Fixes build failure in isolated Docker networks

4. Next.js 15 instrumentation crash fix

  • Removes and
  • These cause on startup

5. Persistent data volume fix

  • Volume mount changed from →
  • Fixes empty database / missing providers on container restart

Testing

✅ Local validation:

  • Container status: healthy (was with HTTP 500)
  • on )
  • V8 heap limit: 6.3 GB (confirmed via )
  • Docker memory limit: 8 GB applied
  • Encryption keys loaded without auth tag failures
  • Database persists with providers/models across restarts

Related

  • Closes: Container OOM/GC storm issues under multi-agent coding load
  • Enables: Headroom proxy () for 60-95% token reduction

diegosouzapw and others added 30 commits July 4, 2026 14:31
…pm scripts (plano testes+CI, Pacote 1) (diegosouzapw#6214)

* perf(test): tsx/esm loader, tsx 4.23 bump, orphan tests recovered, CI runs npm scripts

Pacote 1 (quick wins) do plano mestre testes+CI:

- Swap --import tsx -> --import tsx/esm on the 19 test scripts: the repo is pure
  ESM and the CJS hook costs ~1.3s PER test process (2,462 processes/run).
  Measured: bootstrap 2.3s -> 1.1s; real-suite A/B (tests/unit/db, 12 files)
  22.2s -> 14.1s wall (-36%), 82/82 pass. Non-test scripts keep the full hook.
- Bump tsx ^4.22.3 -> ^4.23.0 (fix for privatenumber/tsx#809 startup regression;
  helps module resolution on big graphs — hook cost unchanged, honest note).
- Recover 22 ORPHAN test files (tests/unit/feature-triage/*.test.mjs, 53 cases,
  53/53 pass) that matched no glob and ran in NO CI job; drop the dead
  'executors' dir from the braces glob.
- Single source of truth for the unit-suite invocation: new test:unit:ci:shard
  (shard via TEST_SHARD env) called by ci.yml test-unit/node24/node26/coverage
  and quality.yml fast-unit — closing two silent drifts: CI was NOT importing
  setupPolyfill.ts, and the fast path glob OMITTED tests/unit/memory + usage.
  quality.yml TIA step + ci.yml test-integration get the tsx/esm swap only.

Validation: full unit suite 21,153 tests / 21,135 pass / 13 skip (5 fails =
known load-flake family, 10/10 green rerun isolated); vitest 237/237; smoke
db+feature-triage 135/135. Record run 2 = ci.yml workflow_dispatch on the
stacked pacote-2 branch (clean runners).

* fix(test): dashboard UI tests keep full tsx hook; recover 15 more .mjs orphans; extend discovery gate to .mjs

Follow-up do dispatch de validacao (run 28720431562), que pegou 2 problemas reais:

1. tests/unit/dashboard/** (11 arquivos, 102 casos) importam componentes React cujo
   grafo puxa @lobehub/icons — o build es/ dele faz require() interno de arquivos com
   sintaxe ESM, que so funciona com o patch CJS do tsx (sem ele: SyntaxError
   'Unexpected token export' no CI; local vira crawl de ~60s/arquivo). Esses 11
   arquivos rodam agora numa 2a invocacao com --import tsx COMPLETO (mesmo shard),
   e o resto da suite mantem tsx/esm (-50% bootstrap). Validado: 102/102.
2. check:test-discovery falhou porque ancorava textualmente os globs nos workflows —
   COLLECTORS atualizado p/ o modelo fonte-unica (ancora = nome do script
   test:unit:ci:shard nos workflows) + varredura ESTENDIDA a .test.mjs, que era o
   ponto cego que deixou os orfaos apodrecerem. A extensao revelou +15 orfaos .mjs
   (top-level + db/) alem dos 22 de feature-triage — TODOS religados via glob
   tests/unit/**/*.test.mjs (171/171 pass). Um deles (encryption-error-handling)
   codificava o contrato PRE-hardening (decrypt falho retornava ciphertext cru —
   vazamento); alinhado ao contrato shipped (null + log) com comentario.

Gate: [test-discovery] OK — 2892 arquivos, 22 collectors, 60 orfaos congelados
(divida rastreada, shrink-only).
…it shards, i18n single job, draft-skip (diegosouzapw#6215)

Pacote 2+3-ci do plano mestre testes+CI (aprovado 2026-07-04). O CI pesado rodava a
suite unit 4x por sync da release-PR (95 jobs, 208 min-maquina) e o ciclo v3.8.44
disparou 123 desses runs (88 cancelados, 0 uteis) porque a release-PR viva fica
aberta o ciclo inteiro.

- D2: matrizes Node 24/26 (build + 8 jobs de teste, ~28% do custo por run) saem do
  per-sync e viram .github/workflows/nightly-compat.yml (diario, fail-fast off,
  resolve a release ativa como o nightly-release-green, abre issue de tracking em
  falha). ci.yml/ci-summary limpos das referencias.
- D3: a matrix Coverage Shard x8 (~18% do custo) e eliminada — o job test-unit roda
  os MESMOS shards sob c8/NODE_V8_COVERAGE e sobe os artifacts coverage-shard-N; o
  job de merge (test-coverage) so repontou needs (padrao do CI do nodejs/node).
  timeout test-unit 15->25min pelo overhead de instrumentacao.
- D4: a matrix i18n de ~40 jobs de <1min (saturava sozinha os 20 slots de
  concorrencia da conta Free) vira 1 job que itera os idiomas com ::group:: por
  idioma e artifact unico com resultados nomeados por idioma (antes 40 result.txt
  colidiam no merge-multiple do ci-summary).
- P3: jobs pesados pulam pull_requests DRAFT (predicado em 10 jobs-raiz; o resto
  pula pela cadeia de needs; ci-summary segue rodando como sinal unico) — a skill
  /generate-release ja abre a release-PR viva como draft e flipa ready no 0a.0a
  (commit eb04fc5 no repo .agents/skills).
- C5 (CodeQL schedule) NAO incluido: bloqueado na acao do dono Settings -> CodeQL
  Default->Advanced (documentado no proprio codeql.yml).

Validacao: js-yaml parse ok; check:workflows zizmor 156 < baseline 159 (ratchet
verde); validacao de execucao = workflow_dispatch deste ci.yml neste branch ate
package-artifact + electron-package-smoke verdes (registrada no PR).
…nt-guard fork-condicional (Pacote 4) (diegosouzapw#6218)

* feat(quality): no-new-warnings per PR via native ESLint bulk suppressions

Pacote 4 do plano mestre testes+CI (aprovado 2026-07-04). O ratchet de
eslintWarnings so rodava no CI pesado (release-PR) -> o drift acumulava invisivel
e explodia na release (+41/+37/+88 por ciclo, rebaselinado as cegas — historico
no proprio quality-baseline.json). Modelo novo (SonarSource Clean-as-You-Code +
ESLint bulk suppressions nativo >=9.24):

- config/quality/eslint-suppressions.json congela a divida existente por
  arquivo+regra: 476 arquivos / 4.273 violacoes.
- npm run lint + lint-staged (pre-commit) + novo job lint-guard no quality.yml
  rodam suppressions-aware: violacao NOVA fica vermelha NO PR que a introduz
  (bulk suppressions ainda eleva estouros de baseline por arquivo a error).
- 3 regras warn promovidas a error em src/** (react-hooks/exhaustive-deps,
  @next/next/no-img-element, import/no-anonymous-default-export) — divida
  existente congelada, ocorrencia nova = erro imediato.
- collect-metrics mede sob o baseline congelado -> a metrica eslintWarnings
  vira 'divida liquida nova' (~0 em regime); baseline apertado 4279->0 no mesmo
  PR (exigencia do require-tighten). Aperto do ESTOQUE congelado: npx eslint .
  --prune-suppressions na reconciliacao da release.
- Principio Zero: lint-guard usa continue-on-error para PR de FORK (report-only;
  a campanha /green-prs aplica o fix via co-autoria) — bloqueante so para
  branches internas, a origem real do drift.

Validacao: negativo (any novo em tests/) exit 1; negativo (img em src/, regra
promovida) exit 1; positivo escopado exit 0; baseline gerado por --suppress-all
no tip (tree inteiro passa por construcao); YAML js-yaml ok.

* fix(quality): clear the 6 residual warnings so lint-guard runs clean at --max-warnings 0

The committed baseline still let 6 warnings through the lint-guard gate:
5 now-unused inline eslint-disable directives (the file-level suppressions
made them redundant — removed via eslint --fix, suppressions regenerated to
absorb the re-exposed occurrences) and 1 anonymous default export in
tests/load/k6-soak.js (outside the src/** severity-override scope — named
the k6 scenario function instead).

Verified on the clean tree: lint-guard exit=0; any-canary (new 'const x: any'
in open-sse) exit=1 — the gate bites on NEW violations while the 4,273
frozen ones stay suppressed (476 files).

* fix(ci): lint-guard continue-on-error must be boolean on non-PR events

github.event.pull_request is undefined on workflow_dispatch — the bare property
expression made the job fail at plan time (run 28722888456: 4 jobs green, run red,
lint-guard never materialized). Guard with event_name check so the expression is
always boolean: PR de fork = report-only (Principio Zero), resto = blocking.
…erhaul (diegosouzapw#6214, diegosouzapw#6215, diegosouzapw#6218)

i18n CHANGELOG mirrors intentionally left to the release reconciliation
(release:sync-changelog-i18n), per cycle practice.
…reeze (base-red) (diegosouzapw#6158)

`src/app/api/oauth/[provider]/[action]/route.ts` grew to 959 lines, past its
frozen cap of 924 (`check:file-size` → Fast Quality Gates red on release/v3.8.44).
The growth came from diegosouzapw#6054 (graceful 400 for keychain-import-only providers / zed):
a doc block, two Sets (KEYCHAIN_IMPORT_ONLY_PROVIDERS, OAUTH_FLOW_ACTIONS) and a
keychainImportOnlyResponse() helper, plus two duplicated guard blocks in GET/POST.

That is a cohesive, self-contained leaf, so extract it to a new
`keychainImportOnly.ts` exposing `keychainImportOnlyGuard(provider, action)`
(returns the 400 NextResponse or null). The two route callsites collapse to a
2-line guard each. route.ts: 959 -> 918 (< 924, freeze restored). No behavior
change.

Tests (Rule #8/#18):
- Existing tests/unit/oauth-keychain-import-only-6041.test.ts (route-level GET/POST
  zed 400) still pass unchanged — behavior preserved.
- New tests/unit/oauth-keychain-import-only-guard.test.ts pins the extracted guard
  in isolation (zed+flow -> 400, normal provider -> null, zed+non-flow -> null).
…ouzapw#31 object toast) (diegosouzapw#6161)

Clicking 'test' on a provider model (e.g. a ClinePass flash model) could freeze
the entire dashboard. Root cause: POST /api/models/test returned an OBJECT in
`error` on the Zod-validation and invalid-JSON paths (`validation.error.format()`
/ a details object). The client does `notify.error(data.error)`, and
NotificationToast renders the message directly as a React child — an object throws
React diegosouzapw#31 ('Objects are not valid as a React child'), crashing the tree = frozen
page instead of a toast.

Fixed in three layers (defense in depth):
1. Server (root cause): /api/models/test now returns a STRING `error` on every
   path — flattens Zod issues to text, returns 'Invalid JSON body' for bad JSON.
2. Client: onTestModel funnels the response through extractApiErrorMessage() so any
   object-shaped error is coerced to a string before notify.error.
3. Toast: NotificationToast coerces title/message via toToastText() — a resilient
   catch-all so no future caller can freeze the page with a non-string.

Tests (Rule #18, both node:test / blocking suite):
- tests/unit/models-test-error-shape.test.ts — asserts STRING error on Zod-fail,
  missing-field, and invalid-JSON (fails on the pre-fix route: 3/3 red -> green).
- tests/unit/notification-toast-coercion.test.ts — toToastText coercion matrix.
… the home page (diegosouzapw#6164)

The blue "Auto-Routing Active — OmniRoute is automatically routing requests
using combo-based strategies" banner was rendered unconditionally on the home
page (`/home`, the default dashboard landing) — it did NOT reflect whether
auto-routing was actually active, and reappeared on every fresh browser / private
window / cleared localStorage (dismissal is stored per-browser). It added noise
to the landing page without conveying live state.

Remove it: drop the <AutoRoutingBanner /> usage + import from home/page.tsx and
delete the now-unused component and its test.
…nly API) (diegosouzapw#6165)

* fix(cline): force upstream streaming for Cline/ClinePass (streaming-only API)

Cline's API (api.cline.bot) only implements streaming (streamText). A
non-streaming request returns HTTP 500 "generateText is not implemented" (Claude
models) or HTTP 502 "empty response" (others). Live-verified on the VPS:
stream:true → works (STREAM_OK), stream:false → fails. This is why testing a Cline
model in the dashboard (the test button sends stream:false) failed.

Fix (reuses the existing isClaudeCodeCompatible mechanism, no new handler):
- Flag `cline` and `clinepass` registry entries with `forceStream: true`.
- In chatCore, OR `providerRequiresStreaming` into `upstreamStream` (line 1591)
  so the upstream request always streams for these providers, while the client's
  original `stream` intent still drives the response format. The existing
  non-streaming branch (parseNonStreamingResponseBody) already accumulates the
  upstream SSE and converts it back to JSON for stream:false clients — the same
  path Claude-Code-compatible providers already use.

Tests (Rule #18): tests/unit/cline-force-stream.test.ts pins the registry flags +
resolveStreamFlag forcing behavior. Live VPS before/after recorded on the PR.

* fix(sse): cline forceStream must stream upstream only, keep client JSON

The diegosouzapw#2081 wiring fed providerRequiresStreaming into resolveStreamFlag,
forcing the client-facing stream flag to true for forceStream providers.
That skips the if(!stream) branch that drains a forced upstream SSE and
converts it back to JSON, so a stream:false caller (model-test button,
plain JSON API) got STREAM_EARLY_EOF instead of a JSON body.

Keep providerRequiresStreaming only on upstreamStream (force upstream to
stream); leave the client-facing stream as the client sent it, so
readNonStreamingResponseBody accumulates the SSE into JSON. The promised
handleForcedSSEToJson (diegosouzapw#2081 comment) was never implemented — this uses
the existing non-streaming SSE-buffering path (same as isClaudeCodeCompatible).

Live-verified on VPS: cline stream:true worked, stream:false failed.
…osouzapw#6170)

* fix(providers): correct Kiro model catalog to real upstream ids

Kiro's API (generateAssistantResponse) returns 400 "Invalid model. Please
select a different model" for any id it does not recognize. The registry
exposed fabricated ids (copied from OmniRoute's own Anthropic catalog) that
Kiro never serves, so every call to them 400'd. Live-verified on the VPS:

  Removed (400 Invalid model):
    - auto-kiro       (no "auto" model id — was sent verbatim upstream)
    - claude-fable-5  (Kiro offers no Fable)
    - claude-opus-4.8/4.7/4.6 (Kiro offers no Opus)
  Corrected:
    - claude-sonnet-4.6 -> claude-sonnet-4.5 (Kiro's Sonnet is 4.5; 4.5 -> 200)
  Kept:
    - claude-sonnet-5 (real Kiro model, plan-gated per account)
    - claude-haiku-4.5, deepseek-3.2, glm-5, minimax-m2.5/m2.1,
      qwen3-coder-next (all proven 200 on the VPS)

Aligns the free-model catalog and drops the orphaned auto-kiro price key.
Regression guard: tests/unit/kiro-catalog-real-models.test.ts (3/3).
Kiro cluster diegosouzapw#6112/diegosouzapw#6113/diegosouzapw#6099.

* test(providers): align stale Kiro-catalog tests to the corrected upstream ids

The fabricated Kiro ids removed in the parent commit (claude-fable-5,
claude-opus-4.8/4.7/4.6, claude-sonnet-4.6) were still asserted as present by
three pre-existing tests, which encoded the bug:
- catalog-updates-v3x: now asserts Kiro does NOT expose Fable 5 / Opus (kept the
  legit cc exposure) and guards the real claude-sonnet-4.5 pricing.
- model-family-fallback-notation: the dot-notation example moves from kiro/ to
  anthropic/ (which genuinely serves Opus/Fable in dot notation) — coverage kept.
- provider-models-route: the Kiro local-catalog assertion now expects the real
  Sonnet 5 / Sonnet 4.5 set and negatively guards the fabricated ids.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
…iegosouzapw#6208)

When ChatGPT Web generates an image as an image_asset_pointer but the pointer
fails to resolve to a downloadable URL (unknown asset scheme, download 403/
expired, oversize), resolveImagePointers returned [] — indistinguishable from
'no image produced' — so the image-generation handler reported the misleading
502 'completed without returning image markdown'. The image genuinely existed
upstream; OmniRoute dropped it silently.

Fix: the executor flags x_image_resolution_failed when a pointer existed but
none resolved (and logs the unresolved asset scheme for follow-up), and the
handler surfaces a truthful 'generated but not retrievable' 502 instead of
'no image markdown'. Adds executorFactory DI for unit testing.

TDD: tests/unit/chatgpt-web-image-silentdrop.test.ts (red -> green), plus the
existing chatgpt-web / image-generation-handler suites stay green.

Reported via community triage (mesh escalated backlog).
…e wiring (diegosouzapw#6211)

* fix(dashboard): providers page data-timeout guard + live-ws standalone wiring

Captura de trabalho em progresso: timeout de dados na página de providers,
ajuste em ProviderLimits e instrumentation-node, com testes novos
(providers-page-data-timeout, live-ws-standalone-wiring).

* chore(quality): rebaseline ProviderLimits/index.tsx file-size (+6, diegosouzapw#6211 data-timeout guard)

Cohesive fix growth from PR diegosouzapw#6211's data-timeout guard on the quota page's two
first-paint fetches (1121->1127). The fast-path PR->release skips check:file-size,
so the bump lands with the PR. Justification recorded in file-size-baseline.json.
…souzapw#6181)

* fix(translator): strip reasoning param for nvidia z-ai/glm-5.2

NVIDIA NIM OpenAI-compatible wrapper rejects the reasoning body field
and returns HTTP 400 "Unsupported parameter(s): `reasoning`".
Add a StripRule scoped to provider=nvidia + model /z-ai\/glm-5\.2/i.
Mirrors PR diegosouzapw#6102 drop pattern (minimax-m2.7 thinking).

* docs(translator): tighten nvidia glm-5.2 strip-rule comment

* fix(translator): anchor glm-5.2 strip rule with word boundary
…tion (diegosouzapw#6177)

NVIDIA NIM API (nvidia/* models) silently truncates the tool list to 128
(the default MAX_TOOLS_LIMIT) because nvidia is not in PROVIDER_TOOL_LIMITS.
Tools beyond index 127 are dropped, causing agents to lose access to
critical tools like task, read, or high-index MCP tools.

Verified that NVIDIA NIM API supports up to 1536 tools by direct testing.
End-to-end confirmed: 198 tools sent, model successfully called tools at
indices 193, 195, and 197 (previously dropped by truncation to 128).

Follows the same pattern as diegosouzapw#5563 (grok-cli: 200), integrated in v3.8.43.
…apw#6200) (diegosouzapw#6209)

* feat(provider): add Claude 5 Sonnet to Claude Web provider (diegosouzapw#6200)

* test(providers): guard claude-web claude-sonnet-5 registry entry (diegosouzapw#6209)

Adds the missing regression test the PR-test-policy gate requires: asserts the
claude-web registry exposes claude-sonnet-5 (Claude 5 Sonnet web) alongside the
existing 4.6 Sonnet / 4.5 Haiku entries. Fails on the release base (no entry).

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

---------

Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
…d address (diegosouzapw#6194) (diegosouzapw#6195)

POSIX shells (bash/zsh) always set HOSTNAME to the machine name. The
.env loader uses first-wins semantics, so HOSTNAME=0.0.0.0 in .env is
silently ignored. This causes the server to bind to the LAN hostname
instead of 0.0.0.0, breaking localhost access and all internal
self-requests (ModelSync, HealthCheck, cloud sync).

The fix compares process.env.HOSTNAME against os.hostname(): when they
match, it's the POSIX auto-set signature and HOSTNAME is ignored.
OMNIROUTE_SERVER_HOST takes precedence as the dedicated escape hatch.

Backward compatibility is preserved: users who set HOSTNAME to a value
that doesn't match the machine name (e.g. Windows CMD/PowerShell users
with HOSTNAME in .env) will still have their value honoured.

Closes diegosouzapw#6194
…ent (diegosouzapw#6213)

Kiro/CodeWhisperer streams Claude's reasoning as native `reasoningContentEvent`
frames when adaptive thinking is enabled, but the Kiro executor had no handler
for them, so `reasoning_effort` requests returned no reasoning. Wire it end to
end:

- translator (openai-to-kiro): enable Kiro thinking when the request carries
  `reasoning_effort`, Anthropic `output_config.effort`, or a `thinking` block
  (`{type:"enabled",budget_tokens}` mapped to a level; `{type:"adaptive"}`
  defaults to `high`, matching Anthropic's documented default). Prepends the
  Kiro `<thinking_mode>`/`<max_thinking_length>` prompt directive and sets
  top-level `additionalModelRequestFields` ({output_config.effort,
  thinking:{type:"adaptive"}, max_tokens}). Gated on `supportsReasoning`; drops
  non-default temperature/top_p (rejected by adaptive-only Claude models).
- executor transformRequest: forward `additionalModelRequestFields` to AWS
  (previously dropped by the strict top-level allowlist).
- executor stream loop: parse `reasoningContentEvent` (and reasoningText
  variants) into the OpenAI reasoning_content channel.

Verified against the live CodeWhisperer stream: reasoningContentEvent frames are
returned, and larger effort/budget measurably deepens reasoning up to the model
cap. Unit tests cover the effort sources, forwarding, temp/top_p stripping, and
native reasoning-frame parsing.
…ation (diegosouzapw#6193)

* fix(chatcore): exempt opencode client from the default 128-tool truncation

The default MAX_TOOLS_LIMIT (128) cap made truncateToolList blind-slice
tools.slice(0, 128), dropping opencode's built-in task tool and part of
its MCP tools when the inbound list exceeded 128 — so models routed
through OmniRoute could not launch subagents or reach all their tools.

Detect the opencode client (any x-opencode-* header, or 'opencode' in
the user-agent) and bypass ONLY the speculative 128 default. A known
provider ceiling (proactive PROVIDER_TOOL_LIMITS or a detected limit)
always wins and still truncates, even for opencode, so upstreams with
real hard limits (e.g. grok-cli 200) keep their 400-avoidance guard.
Non-opencode clients are unchanged.

- requestFormat.ts: add isOpencodeClient(headers, userAgent) + expose it
  on resolveChatCoreRequestFormat.
- toolLimitDetector.ts: add getKnownToolLimit(); getEffectiveToolLimit
  becomes getKnownToolLimit(provider) ?? DEFAULT_LIMIT (byte-identical
  for existing callers).
- upstreamBody.ts: truncateToolList takes bypassDefaultToolLimit and
  encodes the precedence; fix cosmetic debug-log count.
- chatCore.ts: thread the flag into prepareUpstreamBody.
- tests: extend tool-limit-detector unit tests.

* refactor(tools): accept nullable provider in tool-limit resolvers

Address PR review: widen getKnownToolLimit / getEffectiveToolLimit to
(provider: string | null | undefined) to match the call sites in
truncateToolList, and add unit assertions covering null/undefined
providers (getKnownToolLimit -> null, getEffectiveToolLimit -> 128).

---------

Co-authored-by: DKotsyuba <16292493+DKotsyuba@users.noreply.github.com>
Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
* fix(providers): refresh github copilot catalog

Limit GitHub Copilot discovery to the curated supported model set and keep the provider cooldown panel client-safe by moving countdown formatting out of localDb.

* chore(quality): rebaseline providerPageHelpers.ts file-size (+13, diegosouzapw#6154 copilot catalog)

The GitHub Copilot catalog refresh grows the provider-page model-section helper
(1021->1034). Fast-path PR->release skips check:file-size, so the bump lands with
the PR. Justification recorded in file-size-baseline.json.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

---------

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
…ouzapw#6213

The diegosouzapw#6213 kiro adaptive-thinking feature grew openai-to-kiro.ts (853->890) and its
test (1093->1234); the fast-path PR->release does not gate check:file-size on merge,
so the growth accumulated on the release tip. Rebaselined to keep the tip green.
Justification recorded in file-size-baseline.json.
…egosouzapw#6163)

* fix(doctor): resolve two false-positive WARNs (diegosouzapw#6162)

The `omniroute doctor` command reported two warnings on healthy installs
even though the underlying checks actually passed. Both came from the
doctor probing state that already worked; they looked like bugs but users
couldn't tell without manual digging.

Issue 1 — Server liveness HTTP 401
  /api/health and /api/health/degradation both require the management
  token. Doctor called them without auth → 401 → WARN, even when the
  Next.js server was clearly alive and listening.

  Fix: probe the configured health endpoint first; on 401/403, fall
  back to a publicly served static asset (/favicon.ico) to confirm the
  server is alive. WARN now only fires when both probes fail.

Issue 2 — CLI Tools '@/shared' import
  tool-detector.ts (and 3 other cli-helper files) import @/shared/...
  aliases that resolve via tsconfig.json paths. The CLI ships raw TS
  source (no compile step) and runs through tsx, but tsx does not honor
  tsconfig paths at runtime, and tsconfig-paths only hooks CJS
  Module._resolveFilename while doctor uses ESM `import()`.

  Fix: replace @/shared/... with relative imports in the 4 cli-helper
  files. This is the same pattern these files already use for ./config-
  generator/* imports. No new dependency, no architectural change, and
  the fix doesn't regress Next.js itself which keeps using @/shared.

Verified on v3.8.43 (Node v24.17, Windows 11):
  Before: 7 ok, 2 warning(s), 0 failure(s)
  After:  8 ok, N warning(s), 0 failure(s)
    where N accurately reflects which CLI tools are installed and
    configured for OmniRoute (e.g. Hermes Agent installed but not
    pointed at 20128 → 2 real warnings, not 1 false-positive).

Refs diegosouzapw#6162

* fix(doctor): derive fallback URL from primary URL via new URL()

Per Gemini code-assist review feedback: the previous fallback constructed
the /favicon.ico URL from defaults (127.0.0.1:PORT) which ignored custom
host/port/protocol configurations supplied via:
  - OMNIROUTE_DOCTOR_LIVENESS_URL
  - OMNIROUTE_DOCTOR_HOST
  - --liveness-url / --host CLI flags

Parse the primary URL with new URL() to preserve protocol, host, port, and
subpaths. The previous default-based fallback remains as a catch-all for
invalid primary URLs.

* test(doctor): add regression tests for diegosouzapw#6162 fixes

Two new test files lock the fix and satisfy the PR Test Policy gate
("production code change without tests"):

- tests/unit/cli-helper-tool-detector-paths-6162.test.ts
    Locks the @/shared → relative imports fix across all 4 cli-helper
    files. Asserts (a) no @/shared alias remains in the cli-helper
    sources, and (b) each file is importable at runtime via tsx/ESM,
    which would have thrown "Cannot find package '@/shared'" before
    the fix.

- tests/unit/cli-doctor-liveness-fallback-6162.test.ts
    Locks the /favicon.ico fallback in doctor.mjs. Asserts the
    fallback probe exists, derives its URL from the primary URL via
    new URL() (per Gemini review feedback), and that the buggy
    'Server responded with HTTP 401' WARN path is gone.

Both tests use only node:test + node:assert/strict so they slot into
the existing 'test' and 'test:unit' scripts with no extra config.

* test(doctor): fix primary.ok regex in fallback test

The earlier regex /primary\.ok\s*\?/ required a '?' immediately after,
but the actual doctor.mjs code uses a multi-line if-block:

  if (primary.ok) {
    return ok(...);
  }

Use /\bprimary\.ok\b/ instead so the assertion matches the existing
branching.

---------

Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
jmengit and others added 17 commits July 6, 2026 21:16
fix(api-manager): preserve combos in fallback model picker (net +1/-0, test OK). Integrated into release/v3.8.46.
… /v1/models catalog (diegosouzapw#6495)

feat(api): add hidePaidModels setting (net +1/-0, test OK). Integrated into release/v3.8.46.
…6336)

isolate Spark quota + stabilize quota UI (diegosouzapw#6336). Tests 44/44, no file-size drift from its files. Integrated into release/v3.8.46.
launch-codex spawns codex.cmd via shell on Windows (diegosouzapw#6263 pattern). Added resolveCodexSpawn() helper + regression test (2/2), satisfying Hard Rule #18. Integrated into release/v3.8.46.
…ated documentation (diegosouzapw#6349)

add TinyFish web-fetch/search provider + tool (diegosouzapw#6349). Tests green (tinyfish suites + count-guard 170->171). Integrated into release/v3.8.46.
…souzapw#6351)

add GLM team plan quota settings (diegosouzapw#6351). Tests 29/29; reconciled FormData with m365Tier; modal caps bumped for own growth. Integrated into release/v3.8.46.
Codex reset-credit redemption flow (diegosouzapw#6361). Tests 56/56; kept release 'Banked Reset Credits' label (reverted cosmetic rename that broke 2 release tests). Integrated into release/v3.8.46.
…ad of silently falling back (diegosouzapw#6485) (diegosouzapw#6506)

Surface validationErrors for unknown stacked compression engines. Integrated into release/v3.8.46 with a TDD regression.
…egosouzapw#6467) (diegosouzapw#6501)

Intra-message dedup for the session-dedup compression engine (diegosouzapw#6467) + fallbackReason surfacing + fusion rate-limit detail. Integrated into release/v3.8.46 with a TDD regression.
…re-add (diegosouzapw#6499)

Unique default connection name prevents silent overwrite of existing API-key connections. Integrated into release/v3.8.46 with a unit-tested helper; remaining file-size reds are pre-existing base-red drift.
…ai, auto/mimo, auto/gemma, auto/llama, auto/gemini (diegosouzapw#6453) (diegosouzapw#6509)

Provider-family auto combos (diegosouzapw#6453). Integrated into release/v3.8.46; vitest autoCombo suite 11/11 green.
…idePaidModels is on (diegosouzapw#6512) (diegosouzapw#6518)

Exclude paid-only models from the auto/* candidate pool when hidePaidModels is on (diegosouzapw#6512). Integrated into release/v3.8.46; fixed a phantom vitest guard, 4/4 green.
…x duplicated headers and 500 replayed as 200 (diegosouzapw#6408)

The diegosouzapw#6408 request-shape-keyed TTL cache around getUnifiedModelsResponse was not
keyed by DB/settings state, so a write followed by a read within the ~1.5s TTL
replayed the pre-write catalog (dropping newly-eligible models like codex/gpt-5.5
and every specialty catalog routed through it since diegosouzapw#6303). Fold a
modelCatalogCacheVersion (bumped by invalidateDbCache, already called on every
settings/connection/combo/pricing write) into the cache so a state change forces
an immediate miss; merge response headers through a real Headers instance
(fixes X-Request-Id duplication); carry and replay status end-to-end (a mid-build
500 was replayed as 200). Tests call the existing __resetCatalogBuilderRunsForTest
hook in setup, matching v1-models-concurrent-6408.
…w#6366 regression)

diegosouzapw#6366 switched the output base from path.resolve to path.join(process.cwd(),
outputDir) to keep Turbopack's static analyzer from tracing the project root.
path.join mangles an absolute outputDir (a tmp dir in the generator tests) into
cwd/tmp/…, so apply mode reported success while writing nothing at the expected
path. Guard with path.isAbsolute — absolute paths pass through, the relative
production case ('skills') keeps the Turbopack-friendly join form. The existing
agentSkills-generator suite is the regression guard (now 21/21).
… id (CodeQL js/biased-cryptographic-random)

randomNumericId builds a non-secret synthetic device/web id. The v3.8.45 switch
to crypto.getRandomValues closed js/insecure-randomness but 'cryptoByte % 10' is
biased (256 is not a multiple of 10), tripping js/biased-cryptographic-random.
Draw each digit by rejection sampling — discard bytes in the biased tail so the
remaining range divides evenly — giving a uniform distribution (verified) while
staying crypto-backed.
… drift, allowlist diegosouzapw#6303 test consolidation

- Type 12 no-explicit-any errors in 3 new test files (real types, no masking):
  chat-early-schema-validation-6412, models-catalog-envkey-6406, zed-provider.
- Suppress MitmProxyTab.tsx no-html-link-for-pages (false positive: the link is
  an /api/settings/mitm cert download, an API route not a Next page).
- Allowlist tests/integration/v1-contracts-behavior.test.ts net -2 asserts: diegosouzapw#6303
  moved embedding/image shape coverage to models-catalog-route + specialty tests.
- Rebaseline cycle drift measured on the release tip (captain fixes are net-zero,
  verified): cognitive 877->882, cyclomatic 2035->2050, file-size proxies/chat/
  ApiManagerPageClient/ProxyRegistryManager + models-catalog-route testcap +5.
- Fix stale CLAUDE.md: no-explicit-any is error (not warn) in tests/ since diegosouzapw#6218.
Parallel-cycle model (2026-07-04): cut from the frozen release/v3.8.46 tip at the
v3.8.46 release freeze so development continues immediately on v3.8.47 while the
captain closes v3.8.46. Bumps package.json x3 + openapi + lockfile to 3.8.47, adds
the '## [3.8.47] — TBD' living CHANGELOG section, and syncs the 42 i18n mirrors.
Closing v3.8.46 fixes reach this branch via the Phase 5 sync-back.
@oyi77
oyi77 requested a review from diegosouzapw as a code owner July 7, 2026 12:13
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

Gemini encountered an error creating the review. You can try again by commenting /gemini review.

oyi77 added 2 commits July 7, 2026 21:10
… stability

- Dockerfile: Install Python + pipx + headroom-ai CLI for in-container proxy management
- docker-compose.yml: Add resource limits (8GB RAM, 4 CPU), V8 heap 6GB, encryption keys, healthcheck tuning
- src/app/layout.tsx: Switch from Google Fonts to local Inter WOFF2 (fixes build DNS failure)
- Remove Next.js 15 broken instrumentation.ts/ instrumentation-node.ts (fixes startup crash)
- Volume mount: Use ~/.omniroute for persistent data (fixes empty DB issue)

Tested locally: Container healthy with 6GB V8 heap, 8GB RAM limit, 4 CPU cores.
@oyi77
oyi77 force-pushed the feat/headroom-resource-optimizations branch from 35117ba to 5610bfd Compare July 7, 2026 14:10
…endpoint

- Add HEALTHCHECK_BATCH_SIZE (default 20) to cap refreshes per tick
- Skip 3s stagger delay between rotating-provider refreshes (kiro, etc.)
- Move sweep lock guard earlier to prevent overlapping sweeps under load
- Add getProviderConnectionsHealth() that skips decryption of 2113 rows
- Health route now uses lightweight query: 0.08s vs 28s on cold endpoint
@oyi77 oyi77 closed this Jul 7, 2026
@oyi77
oyi77 deleted the feat/headroom-resource-optimizations branch July 7, 2026 17:40
diegosouzapw added a commit to oyi77/OmniRoute that referenced this pull request Jul 12, 2026
Add checkConnectionCapacity guard with 429 + Retry-After in handleChat().
Introduce OMNI_MAX_CONCURRENT_CONNECTIONS env-bound cap, disabled (0) by
default so existing deployments are unaffected until an operator opts in.

Reconstructed from PR diegosouzapw#6590, isolating only the backpressure change —
the original branch also carried unrelated headroom/docker/perf work
from the author's separate diegosouzapw#6572 branch.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
diegosouzapw added a commit that referenced this pull request Jul 12, 2026
Add checkConnectionCapacity guard with 429 + Retry-After in handleChat().
Introduce OMNI_MAX_CONCURRENT_CONNECTIONS env-bound cap, disabled (0) by
default so existing deployments are unaffected until an operator opts in.

Reconstructed from PR #6590, isolating only the backpressure change —
the original branch also carried unrelated headroom/docker/perf work
from the author's separate #6572 branch.

Co-authored-by: oyi77 <oyi77@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
…#6590)

Add checkConnectionCapacity guard with 429 + Retry-After in handleChat().
Introduce OMNI_MAX_CONCURRENT_CONNECTIONS env-bound cap, disabled (0) by
default so existing deployments are unaffected until an operator opts in.

Reconstructed from PR diegosouzapw#6590, isolating only the backpressure change —
the original branch also carried unrelated headroom/docker/perf work
from the author's separate diegosouzapw#6572 branch.

Co-authored-by: oyi77 <oyi77@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…#6590)

Add checkConnectionCapacity guard with 429 + Retry-After in handleChat().
Introduce OMNI_MAX_CONCURRENT_CONNECTIONS env-bound cap, disabled (0) by
default so existing deployments are unaffected until an operator opts in.

Reconstructed from PR diegosouzapw#6590, isolating only the backpressure change —
the original branch also carried unrelated headroom/docker/perf work
from the author's separate diegosouzapw#6572 branch.

Co-authored-by: oyi77 <oyi77@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.