Skip to content

fix(resilience): stop unbounded queue that hangs 6min until Aborted - #12715

Merged
diegosouzapw merged 27 commits into
diegosouzapw:release/v3.8.51from
maxmad64bis:feat/cascade-pr1-resilience
Sep 10, 2026
Merged

diegosouzapw merged 27 commits into
diegosouzapw:release/v3.8.51from
maxmad64bis:feat/cascade-pr1-resilience

Conversation

@maxmad64bis

@maxmad64bis maxmad64bis commented Sep 4, 2026 •

Copy link
Copy Markdown
Contributor

⚠️ base-red inherited: #12732

Summary

Related Issues

Validation

  • Change type: other
  • Focused tests and category gates from the golden path
  • npm run lint — targeted eslint on the touched files is clean; the full run is red on the base (🔴 Release branch not green: release/v3.8.51 #12732), so the box stays unchecked.
  • Reconciled with the current active release base; focused checks rerun afterward
  • Production-code changes include a new or updated automated test in this PR

Focused suites, all green on this branch after the change: rate-limit-remaining-budget (5/5), rate-limit-manager-queue-bound (4/4), rate-limit-execution-per-conn (4/4), plus the rate-limit-execution-timeout-message-4165 + provider-rate-limit-overrides-schema companions (25/25). No VPS gate applies — it's local queue/budget plumbing exercised by unit tests, no upstream round-trip involved.

Tests Added Or Updated

  • tests/unit/rate-limit-remaining-budget.test.ts — shared budget gate+slot+queue (remaining propagates, 503 on exhaust).
  • tests/unit/rate-limit-manager-queue-bound.test.ts — bounded Bottleneck queue with anti-orphan guard remaining<=0 not queued, abort vs timeout distinct).
  • tests/unit/rate-limit-execution-per-conn.test.ts — per-connection executionMaxWaitMs and upstream clamp.
  • Focused suite 13/13 across the 3 new files pass, companions 25/25 green.

Coverage Notes

Touches open-sse/handlers/chatCore.ts, open-sse/handlers/chatCore/queueBudget.ts (new), config/quality/file-size-baseline.json, open-sse/services/rateLimitManager.ts, open-sse/services/providerDefaultRateLimit.ts, src/lib/resilience/settings/types.ts, src/shared/validation/schemas/provider.ts.
Covered by the 3 new suites plus rate-limit-execution-timeout-message-4165.test.ts and provider-rate-limit-overrides-schema.test.ts. No file loses coverage.

Reviewer Notes

  • Queue time (503 on budget exhaust) vs queue depth (429 on SEMAPHORE_QUEUE_FULL; 504 on execution expiration) intentionally distinct — only bare time-budget timeouts are shaped to legacy 503.
  • executionMaxWaitMs is max(perConn, upstream) via opts:{executor,providerSpecificData} passed from handleChatCore to withRateLimit.
  • chatCore.ts is a frozen file already at its ceiling. The admission error shaping lives in open-sse/handlers/chatCore/queueBudget.ts; the remaining call-site wiring (+15 lines) is rebaselined in config/quality/file-size-baseline.json with a note, like the existing _rebaseline_* entries. The gate-ordering contract chatcore-hierarchical-admission passes its first case; its second case fails the same way on the base tip.

@maxmad64bis
maxmad64bis force-pushed the feat/cascade-pr1-resilience branch from e44c35d to f9c7e60 Compare September 4, 2026 08:12
@maxmad64bis
maxmad64bis marked this pull request as ready for review September 4, 2026 08:50
@maxmad64bis
maxmad64bis force-pushed the feat/cascade-pr1-resilience branch 2 times, most recently from ac18315 to 250c39a Compare September 4, 2026 08:55
Per-connection maxWaitMs budget is now shared across acquisition gates, provider-default slot wait, and Bottleneck queue (fail-closed 503 on exhaust). Prevents the 45s gate+slot+queue sum that previously hung requests until manual abort. Execution backstop executionMaxWaitMs is now overridable per connection and clamped to the real upstream fetch timeout so it never kills a mid-flight request.
@maxmad64bis
maxmad64bis force-pushed the feat/cascade-pr1-resilience branch from 250c39a to 2061fd2 Compare September 10, 2026 11:51
diegosouzapw and others added 22 commits September 10, 2026 09:11
Contagem de migrations 171 → 172 após a `175_call_logs_provider_stats_indexes.sql` do diegosouzapw#12832. Medido com `ls src/lib/db/migrations/*.sql | wc -l`.

Falha minha de processo: depois da onda 1 desta leva eu medi file-size, api-typecheck, changelog-integrity e colisão de migration — não o `check:docs-counts`. O drift ficou vivo até a onda 2 esbarrar nele.

41 mirrors de `llm.txt` regenerados pelo script do projeto. Aprovado por você para tocar `AGENTS.md`, mesma classe do diegosouzapw#12970.
…thout a documented free tier (diegosouzapw#12744)

Validado numa worktree combinada com a onda de roteamento/free-tier desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (288), check-file-size OK após rebaseline, 84/85 nos testes focados no runner Node.

Parar de honrar `isFree:true`, `:free` e `0/0` vindos do upstream **antes** de consultar o catálogo é a inversão certa: hoje um provider sem tier livre documentado consegue se declarar grátis e o listing diverge do roteador `auto/*`. Checar as heurísticas depois do hit de catálogo fecha a porta sem quebrar o caminho de linhas custom locais, que continuam confiáveis pelo caminho próprio.

O `isFreeModel("or", …)` com um alias que não existe é o tipo de bug que passa despercebido porque falha silenciosamente para o lado permissivo.

**Integração:** o `decideHidePaid` que o diegosouzapw#12795 extraiu passou a usar o seu `isFreeForProvider` por id, em vez de OR-ear os dois aliases num único `freeProvider`. A forma por id é a garantia que esta PR estabelece, então ela prevaleceu.
…nifest (diegosouzapw#12786)

Validado numa worktree combinada com a onda de roteamento/free-tier desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (288), check-file-size OK após rebaseline, 84/85 nos testes focados no runner Node.

Metadado de display derivado da mesma fonte da decisão, com um gate STRICT que quebra o CI se a contagem do manifesto divergir do catálogo — é o detalhe que impede a tag de virar mentira daqui a três meses.

Registrar que 77 entradas de catálogo viram 76 no manifesto porque o `arcee-ai` ainda não tem entrada no registry é exatamente o tipo de discrepância que costuma virar bug fantasma.
…ty/breaker (diegosouzapw#12794)

Validado numa worktree combinada com a onda de roteamento/free-tier desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (288), check-file-size OK após rebaseline, 84/85 nos testes focados.

Tratar breaker aberto como equivalente a fechado no scoring de snapshot é pior que não pontuar: afirma saúde onde há falha conhecida. Trocar as três constantes neutras por valor observado é a correção, e manter preço e orçamento fora do escopo mantém a PR revisável.

**Integração:** o `computeSnapshotWeights` conflitou com o diegosouzapw#12731, que adiciona peso de `reliability` a partir de `failureRate`/`errorRate`. Os dois cobrem chaves diferentes e compõem — ficaram ambos: reliability do diegosouzapw#12731, health via breaker e quality desta PR, quota neutro nos dois. Nenhum dos dois lados foi descartado.
…diegosouzapw#12731)

Validado numa worktree combinada com a onda de roteamento/free-tier desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (288), check-file-size OK após rebaseline, 84/85 nos testes focados.

`reliability-first` que não pesava reliability é o defeito mais constrangedor possível num mode pack, e a causa é clara: `modePacks.ts:13` substituía os defaults por inteiro. Financiar os novos pesos com `quota`/`costInv`/`tierPriority` mantendo cada pack somando 1.0 é a parte que exige cuidado e você fez.

Manter `quality-first` em 0.03, igual ao default, para que ele não fique mais fraco que `balanced`, é o tipo de detalhe que só aparece quando se checa a tabela inteira.

**Integração:** conflitou com o diegosouzapw#12794 no `computeSnapshotWeights`; os dois compõem e ambos ficaram.
…iegosouzapw#12790)

Validado numa worktree combinada com a onda de roteamento/free-tier desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (288), check-file-size OK após rebaseline, 84/85 nos testes focados.

Duas cópias da mesma ordem de provider com um "keep in sync" implícito é dívida que cobra juros a cada provider novo. Uma definição com re-export nos dois lados resolve a classe.

O `xao/*` ordenando depois de todo provider conhecido em vez de junto do `xai-oauth` é um sintoma concreto de que a duplicação já estava divergindo.

**Integração:** `scripts/quality/run-all-gates.mjs` conflitou com o `check:pricing-freshness` que entrou pelo diegosouzapw#12792 na mesma onda. Aditivo — os dois gates coexistem.
…s, pooled latency bootstrap, fresh tier cache (diegosouzapw#12792)

Um modelo grátis fora da tabela herdando $5/$15 por milhão e afundando no roteamento cost-aware é o defeito mais caro desta onda: silencioso, e inverte exatamente a decisão que o operador quer.

Parar de chutar 1500ms de latência para modelo desconhecido e usar a mediana observada do pool — com contador de quantas vezes o chute dispara — é trocar heurística por medição do jeito certo. O contador é o que permite saber se valeu.

Revalidei após reconstruir a branch sobre o tip: **33/33** nas suítes da PR, typecheck:core limpo, `check-api-typecheck` OK (289).

**Duas integrações:**

1. `computeSnapshotWeights` conflitou com o diegosouzapw#12794 (health via breaker + quality), já mergeado. Os dois compõem e ambos ficaram: o seu termo de `reliability` — que era a única chave que o caminho de snapshot ainda ignorava — mais o health observado e o quality do diegosouzapw#12794.
2. `scripts/quality/run-all-gates.mjs` conflitou com o `check:provider-order-sync` do diegosouzapw#12790. Aditivo, os dois gates coexistem.

**Nota de dívida:** o `virtualFactory.ts` cruzou o teto de 1200 linhas pela primeira vez (1187 → 1207) somando esta onda. Congelei em vez de dividir e registrei os dois candidatos a extração na justificativa — `computeSnapshotWeights` (~85 linhas) e o grupo de elegibilidade de credencial (~70). Qualquer um dos dois volta o arquivo para baixo do cap.
…op reasons (diegosouzapw#12795)

Um pool `auto/*` vazio que não diz por que está vazio é a pior forma de falha: o operador vê ausência e não sabe se é config, cota ou catálogo. Registrar qual estágio removeu quantos, e carregar `dropReason` em cada entrada de `quotaHealth.providers`, transforma silêncio em diagnóstico.

Os três defeitos achados de carona valem tanto quanto a feature — em especial a cota livre recorrente sem teto sendo tratada como desconhecida em vez de segura, que é justamente o caso que mais aparece.

Revalidei após reconstruir sobre o tip: **15/15**, typecheck:core limpo, `check-api-typecheck` OK (289).

**Uma mudança minha no `catalogPaidFilter.ts`.** Mantive a sua extração — ela é mais limpa que o predicado inline — mas troquei o corpo para usar `isFreeForProvider` por id. O módulo OR-eava `providerHasFreeModels(resolved) || providerHasFreeModels(canonical)` num único `freeProvider` e depois testava `isFreeModel` em cada um. Com isso, um id cujo próprio provider não documenta tier livre passa a ler como grátis sempre que o alias irmão documenta — que é exatamente o buraco que o diegosouzapw#12744 fechou e já está no tip. Por id preserva a garantia; nenhuma outra linha do módulo mudou.

**Dívida registrada:** o `virtualFactory.ts` foi de 1187 para 1219 somando esta onda e cruzou o teto de 1200 pela primeira vez. Congelado com os candidatos a extração nomeados na justificativa.
…diegosouzapw#13044)

Batch 1 of the locale expansion: Greek, Croatian, Serbian, Lithuanian, Estonian, Latvian, Slovenian, Maltese and Irish across the dashboard catalog, docs mirrors, CLI catalog, README, locale index and the site. 42 → 51 locales.

Also fixes the ICU literal escape the translation backend dropped around angle placeholders, four translations that invented or renamed a placeholder, the language bars that linked to mirrors that do not exist, and the migration count drift (171 → 172).

⚠️ base-red inherited: diegosouzapw#12732 — the four unit shards and Fast Quality Gates fail identically on unrelated PRs cut from the same base.
…e History tab (2.9) (diegosouzapw#12677)

* feat(dashboard): pure model to compare two orchestration runs

* feat(dashboard): compare-mode selection in the History grid

* feat(dashboard): side-by-side comparison panel in the History tab (2.9)

* chore(dashboard): compare-runs i18n + changelog

* fix(dashboard): compare-panel loading state, height bound, delta legend, ARIA level

Final-review fix wave for PR-A (Orchestration Canvas Fase 3):

- Give each compare-panel side an explicit fetch status (loading/ok/error) so the
  Events metrics row shows "—" instead of a misleading real "0"/delta while a side
  is still loading or after its fetch failed.
- Bound the compare panel's height (max-h-[45vh], overflow-y-auto, shrink-0) so a
  run with many activities can no longer collapse the History grid to zero height.
- Clear the compare selection when the History preset changes, since a stale pick
  can fall outside the new range.
- Add a delta-column legend (new compareDeltaLegend i18n key, translated into all
  41 non-English locales) so operators know which side a positive delta favors.
- Use role="status" (not role="alert") for the informational compareDifferentIdentity
  banner, reserving role="alert" for actual per-side fetch failures.
- Restore ro.json's compareCost to the true cognate "Cost" (was distorted to
  "Cheltuieli" to dodge a byte-identical-to-English heuristic); audited the other
  40 locales for the same pattern across the 9 compare* keys, no other instance found.

* refactor(dashboard): split compare-runs/history-tab functions to clear complexity ratchets

CompareRunsPanel (92 lines, max-lines-per-function) and HistoryTab (85 lines,
same rule) exceeded the 80-line function cap; compareRuns.ts's a2aEventsFrom
exceeded the cognitive-complexity cap (16 > 15). Extract pure/presentational
helpers (ComparePanelHeaderBar, SideErrorRow, ComparisonMetrics, a2aEventFrom,
HistoryStatusRows, refreshNowMsOnActionDone) with no behavior, DOM, i18n or
aria change.

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
…zapw#12392) (diegosouzapw#12983)

* fix(dashboard): keep the first failure timestamp in sourceStale

buildSourceStatuses stamped nowIso on every failing source at every poll, so
the stale indicator reported "since the last poll" instead of the first
failure — and, because snapshotContentKey serializes sources, the snapshot
identity churned on every tick while any source was down.

The failing branches now reuse the staleSince already held by that source in
the previous status list, via the functional setStatuses updater (no ref read
during render, no setState inside an effect body).

Refs diegosouzapw#12392

* fix(dashboard): flag a source that starts failing after it had data

buildRootAndSourceEdges only materialized a placeholder SourceNode when the
failing source had no node at all. A source that already had work nodes and
then started failing (or went offline) kept its healthy-looking SourceNode
forever: no ⚠, no stale styling, no `sourceStale` line — the operator saw a
normal source while it was actually broken.

Now every non-ok/offline source is flagged: when its SourceNode is missing the
placeholder is created as before; when it exists, the node is replaced by a
copy carrying `sourceIssue` and `staleSince`. The copy (never a mutation) keeps
the function pure — the original object is still referenced by the caller's
`parts`, the same trap the droppedByState aliasing fix covered.

Tests: three cases in tests/unit/ui/orchestrationModel.test.ts — existing node
starting to fail (flags set, work nodes kept, no duplicate node, input object
untouched), existing node going offline (no invented staleSince), and a healthy
source staying free of both fields.

Refs diegosouzapw#12392

* fix(dashboard): canvas polish batch (diegosouzapw#12392)

Seven pointwise fixes on the Orchestration Canvas, each covered by a test:

1. Debounce x chip race: every chip/clear write in OrchestrationToolbar now cancels
   the pending search timer first. Left armed, it fired ~300ms later with a setParams
   closed over the pre-chip query string and silently reverted the chip.
2. The search input carries an aria-label (searchPlaceholder) — the placeholder alone
   is not an accessible name.
3. parseCsvSet trims each token, so `?state=running, failed` parses like the unpadded
   form instead of dropping the padded value.
4. toggleCsv was duplicated in the toolbar and the page client; both now import the
   single definition from the new model/urlParams.ts (pure, never mutates its inputs).
5. AgentsTab tells "nothing running" apart from "the filter matched nothing": with an
   active filter and no work node it renders noMatches + a clear-filters button instead
   of the setup CTAs, which would be wrong advice there.
6. Particle cap: orchestrationToFlow stamps `particles` on every edge and turns it off
   above PARTICLE_EDGE_CAP (40) simultaneously active edges — StatusEdge then renders
   the colored stroke without its 3 SMIL particles per edge.
7. The drawer's error banner clears when an action succeeds, so a recovered failure
   does not stay on screen.

Only `noMatches` is added to en.json here; the other locales are task B4.

Refs diegosouzapw#12392

* chore(dashboard): canvas polish i18n + changelog

Real translations for orchestration.noMatches in the 41 non-English locales,
each one written against that file's own neighbouring keys (emptyTitle,
stateRunning, searchPlaceholder) so the wording for "task" and "filter"
matches what the locale already uses. No i18n:sync-ui, no __MISSING__ left.

Adds the changelog fragment for the nine PR-B fixes.

Closes diegosouzapw#12392
…ck (diegosouzapw#12880)

Validado numa worktree combinada com a onda de streaming desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 88/88 nos testes focados.

`agent_message` chegando num fallback de Chat Completions é um item que o cliente não sabe interpretar; mapear ou descartar é a escolha certa, e escolher por item em vez de derrubar a resposta inteira mantém o fallback útil.
…ream stays silent (diegosouzapw#12828)

Validado numa worktree combinada com a onda de streaming desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 88/88 nos testes focados.

O diegosouzapw#12151 cobriu só metade: passthrough emitia o chunk final de usage, translate calculava a estimativa **depois** de fechar o stream, então o número só chegava ao log do servidor e nunca ao cliente. Fechar essa metade é o que faz a feature existir de fato.

Não emitir segundo chunk quando o upstream já mandou usage real é o detalhe que impede a correção de virar contagem dobrada.
…p-dropped event array (diegosouzapw#12718)

Validado numa worktree combinada com a onda de streaming desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 88/88 nos testes focados.

Reconstruir o resumo a partir de um array que o próprio coletor já truncou por cap produz um resumo que parece completo e não é — pior que resumo ausente, porque não se distingue. Parar de reconstruir dali é a correção.
…ycle/heartbeat events, never real content (diegosouzapw#12741)

Validado numa worktree combinada com a onda de streaming desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 88/88 nos testes focados.

Um stream que só emite eventos de ciclo de vida e heartbeat, sem conteúdo nenhum, é falha disfarçada de sucesso: o cliente espera até o timeout dele. Falhar rápido devolve o controle.
…2941)

Validado numa worktree combinada com a onda de streaming desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 88/88 nos testes focados.

Estender a rotação que já existe para 429 ao 403 de bloqueio geográfico é a generalização certa, e manter a rejeição de fingerprint (Cloudflare 1010) fora dela é o que impede a rotação de queimar todas as contas contra uma recusa que não é de egresso.

Nota: os checkboxes de validação do corpo ficaram em branco, mas o diff traz dois arquivos de teste — vale marcar da próxima para o revisor não precisar conferir.
…n in-memory pending store (diegosouzapw#12854)

Diagnóstico por captura de pacote em tráfego real, com o `400 previous_response_not_found` reassemblado do tcpdump três vezes no mesmo loop de tool-calling — isso é evidência, não hipótese. A causa é limpa: `detail_state` só vira `'ready'` depois de uma escrita fire-and-forget enfileirada num worker único, e o cliente já tem o id de resposta antes disso. Semear a ponte **antes do primeiro await** é o que faz a correção não custar latência.

Revalidei sobre o tip: **19/19**, typecheck:core limpo, check-file-size OK.

**Estava draft e eu marquei como ready.** Não havia gate declarado — nem RFC pendente, nem decisão de produto em aberto — e passou na validação; a diretiva permanente do dono para esta campanha é avaliar draft como qualquer PR e promover quando passa. Se a intenção era segurar por outro motivo, me avise que eu reverto.

**Um conserto meu na sua branch.** O `typecheck:core` falhava com `TS2345` em `callLogs.ts:489` — e falhava **na sua branch sozinha**, não por interação com a onda; confirmei isolando. O call site fazia cast para `{ clientRawRequest?: unknown; clientResponse?: unknown }`, mais frouxo que o `ContinuationPipeline` que o parâmetro exige, e `unknown` não assina para os membros tipados. Exportei o `ContinuationPipeline` do próprio store e usei ele no cast, em vez de alargar o tipo do parâmetro: o contrato passa a ter um nome só, no lugar onde ele já vivia.

**Sobre a sua Reviewer Note do Map sem limite de contagem:** concordo que vale registrar. Entradas pequenas com TTL de 60s auto-expirando não justificam sizing agora, mas se aparecer burst sustentado o sintoma será memória, não erro — e aí a nota está aqui.

Também carreguei o rebaseline de `chatCore.ts` (6021→6026) e `stream.ts` (3080→3098), que a onda de streaming inteira faz crescer.
…nal (diegosouzapw#12639) (diegosouzapw#12988)

* feat(api): hydrate memoryHits from the persisted history event

`GET /api/a2a/tasks/[id]` falls back to the persisted history row once a task
leaves the in-memory TTL window, and `reconstituteHistoricalTask` hard-coded
`metadata: {}` — so the drawer's "Memory used" section vanished for any
historical task, even though `executeA2ATaskWithState` had already written a
`memory_hits` event with the hits.

The fallback now reads that event: `data_json` is parsed and, when it yields at
least one well-formed hit, exposed as `metadata.memoryHits`. The event itself is
filtered out of `events` — it is observability, not a state transition, and
without the filter it leaked into the timeline as a duplicate of the row's
current state.

Reading is defensive throughout, mirroring `DrawerMemory`'s own validation: the
payload is caller-influenced and unvalidated end to end, so `JSON.parse` runs
inside `safeJsonParse`, non-arrays are rejected, and each entry must carry `id`,
`key`, `type` and `snippet` as strings (a non-string field would be rendered as
a React child and take the drawer down). Malformed input degrades to
`metadata: {}` and a 200 — never a 500.

Refs diegosouzapw#12639

* fix(a2a): bound the memory recall with its own deadline

collectMemoryHits() runs BEFORE the skill handler and had no deadline at all,
so a slow memory backend delayed the start of every A2A task — the HTTP
genericBackend alone defaults to a 30s timeout.

The search now races a MEMORY_RECALL_TIMEOUT_MS (1500ms) deadline. Overshooting
degrades exactly like any other recall failure: empty hits, a warn log, and the
task proceeds normally (best-effort contract unchanged, nothing propagates).
The deadline timer is cleared in a finally on BOTH paths so no handle is left
holding the event loop open, and MemoryHitsDeps.timeoutMs makes it injectable
so the tests cost milliseconds instead of 1.5s of wall clock.

Refs diegosouzapw#12639

* fix(dashboard): carry conductor requirements and focus the repeated task

The drawer's "Repeat" for a Conductor task dropped the runner/model pinning and
left the operator staring at the finished run:

- `hubTaskSchema` now parses the hub's `requirements` (`.catch(null)` so an odd
  shape never fails the whole task parse), and `ConductorTaskDetail` exposes
  `cli`/`model` (`null` when the hub sends none).
- `repeatReqForConductor` carries `cli`/`model` when present and OMITS them
  otherwise — the route's Zod takes both as optional strings, so a `null` would
  400. The two fields are independent.
- `performAction` reads the response body once and returns it, so the repeat can
  report the CANVAS id of the created task (`task_id` / `data.id` /
  `result.task.id`, each with its node prefix). `OrchestrationPageClient` then
  refetches and focuses it via `?node=`; History keeps its current behavior.
- `conductor-routes-auth.test.ts` covers the creation route through its `ROUTES`
  array; the duplicated source assertion left `conductor-create-route.test.ts`.

Refs diegosouzapw#12639

* chore(a2a): follow-ups changelog

Changelog fragment for the five items PR-C delivers from diegosouzapw#12639.

The sixth item on the issue — an authenticated panel path for A2A task
creation — stays deliberately out of scope and is recorded as such in a
comment on the issue rather than silently dropped: the JSON-RPC endpoint
accepts API keys only, and widening that endpoint's auth surface to serve a
UI convenience is the operator's call, not the implementation's.

Closes diegosouzapw#12639
… across flow surfaces (diegosouzapw#12378) (diegosouzapw#13203)

* refactor(ui): move shared flow colors to the orchestration status tokens

FLOW_EDGE_COLORS and TokenHealthBadge were pinned to the fixed dark-mode
hexes in STATUS_HEX, so both rendered dark-theme green/amber/red on a light
background. They now read the theme-aware --orch-status-{success,warning,
error,muted} custom properties introduced in Fase 2. The dark values of
those tokens are exactly the old hexes, so dark mode is unchanged and only
light mode gains contrast. `idle` was already a CSS var, which is the
precedent proving a var() resolves in a ReactFlow edge stroke.

Five call-sites built translucent variants by concatenating an 8-bit alpha
suffix onto the palette hex (`${FLOW_EDGE_COLORS.error}40`), which cannot
work with a var(). They move to a new documented helper, flowColorAlpha(),
that wraps color-mix() — the same approach orchStateBadgeBg() already uses
in the orchestration model. Percentages mirror the old suffixes
(20 -> 13%, 30 -> 19%, 40 -> 25%).

STATUS_HEX stays exported as the dark-mode mirror; it now has no production
consumer. globals.css needed no change — all five tokens already existed in
both themes.

The colour assertions in the topology, combo-live and design-grid suites
were aligned to the tokens, never weakened: every hex equality became an
equality against the corresponding var(). design-grid additionally now
asserts each token is defined in BOTH themes.

Refs diegosouzapw#12378

* refactor(ui): finish the status-token migration across flow surfaces

Sweeps the five state hexes across the remaining flow surfaces, following D1:

- ComboLiveStudio: active/error provider pills and the run-outcome tri-state.
- CompressionCockpit / WaterfallInspector / IoNode: the savings readouts and the
  savings quality ramp (>=30 success, >=15 warning, else muted).
- EngineNode: the same ramp, plus the running state, whose glow moved to
  flowColorAlpha — the literal #f59e0b40 suffix is invalid once the value is a var().
- WaterfallInspector: a skipped step now reads as muted rather than a bare grey hex.

Deliberately NOT migrated, because they are categorical or brand palettes rather
than state: STRATEGY_COLORS (routing-strategy hues), LAYER_COLORS (compression
layer pills), the provider brand color in ProviderTopology, and IoNode's
indigo/green input-output identity pair. The new test asserts both halves — what
became a token AND what stays hex — so a later sweep cannot silently swallow a
categorical palette.

Refs diegosouzapw#12378
…egosouzapw#12876)

Validado numa worktree combinada com a onda de dashboard/monitoring desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 130/131 nos testes focados — a falha restante é asserção de tempo de parede sob carga, verde 6/6 isolada.

Um health que diz "falhou" sem dizer **qual** conexão obriga o operador a cruzar logs para achar o óbvio. Expor os ids das que falharam é o que transforma o endpoint em ferramenta de diagnóstico.
…iegosouzapw#12882)

Validado numa worktree combinada com a onda de dashboard/monitoring desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 130/131 nos testes focados — a falha restante é asserção de tempo de parede sob carga, verde 6/6 isolada.

Paginar e completar o stream em `GET /v1/models` é a correção certa para catálogo grande: um payload único que cresce com o número de providers vira timeout silencioso no cliente, não erro.
…n on one model 402 (diegosouzapw#12875)

Validado numa worktree combinada com a onda de dashboard/monitoring desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 130/131 nos testes focados — a falha restante é asserção de tempo de parede sob carga, verde 6/6 isolada.

Envenenar a conexão inteira por um 402 de **um** modelo é o erro clássico de granularidade em provider openai-compatible com múltiplos upstreams — derruba modelos que estavam saudáveis. Restringir ao modelo afetado é o comportamento correto, e é a mesma distinção que o guia de resiliência faz entre cooldown de conexão e lockout de modelo.
hartmark and others added 4 commits September 10, 2026 10:42
…instead of "(empty)" (diegosouzapw#12727)

Validado numa worktree combinada com a onda de dashboard/monitoring desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 130/131 nos testes focados — a falha restante é asserção de tempo de parede sob carga, verde 6/6 isolada.

"(empty)" para um nó de ferramenta ainda não resolvido é informação errada, não ausência de informação — o usuário lê como "não retornou nada". Spinner de pendente diz a verdade.
…egosouzapw#12937)

Validado numa worktree combinada com a onda de dashboard/monitoring desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 130/131 nos testes focados — a falha restante é asserção de tempo de parede sob carga, verde 6/6 isolada.

Sentinela que não se explica (`—`, `?`) faz o leitor inventar a razão. Explicar no hover é metade; o `check:radar-sentinels` é a outra — sem o gate, a explicação apodrece na primeira coluna nova.
…iegosouzapw#12884)

Um `ALL_TARGETS_SKIPPED` 503 que não diz qual janela esgotou é opaco justamente no momento em que o operador mais precisa saber. Alinhar os rótulos de janela AUTH com os da API de uso fecha a outra metade: dois nomes para a mesma coisa fazem o dashboard e o erro parecerem discordar.

Revalidei sobre o tip: **6/6**, typecheck:core limpo, check-file-size OK.

**Dois consertos meus na sua branch.**

1. `typecheck:core` falhava com `TS2345` em `comboAttemptLoop.ts` (linhas 130 e 416): o `QuotaSkipTarget` declarava `connectionId?: string`, mas o `ResolvedComboTarget` carrega `string | null` para alvo não-pinado. Alarguei para `string | null` no tipo de diagnóstico em vez de estreitar o call site — o módulo só **lê** o campo e a linha 29 já narrowa com `typeof === "string"`, então null não custa nada ali. Isso apareceu porque o `comboAttemptLoop` mudou de forma no diegosouzapw#12746/diegosouzapw#12811, mergeados nesta mesma campanha depois que você cortou a branch.

2. O `roundRobinCombo.ts` foi de 1198 para 1205 e cruzou o teto de 1200 para arquivo novo. Congelei com justificativa: o arquivo já nasceu em 1198 quando o diegosouzapw#12811 o levantou de dentro do `combo.ts`, e os diagnósticos em si vivem no `quotaSkipDiagnostics.ts`, sob o cap. Registrei que a próxima extração natural é o corpo do attempt loop, mas que ele acabou de ser movido e deve assentar antes de ser cortado de novo.
@diegosouzapw
diegosouzapw merged commit a152eb9 into diegosouzapw:release/v3.8.51 Sep 10, 2026
4 of 7 checks passed
thinh0704hcm added a commit to thinh0704hcm/OmniRoute that referenced this pull request Sep 12, 2026
Brings canonical up to upstream release/v3.8.51 (51 commits), including the
queue-bound resilience fix a152eb9 (diegosouzapw#12715).

Conflicts resolved:
- open-sse/services/rateLimitManager.ts: took upstream's per-connection queue
  budget (maxWaitMs shared across the provider gate, default slot and the
  Bottleneck queue) and re-applied the fork overlay's linked abort signal on top,
  so an expired execution backstop also aborts the in-flight upstream request
  instead of leaking it and holding a slot queued work must wait on.
- open-sse/utils/stream.ts: took upstream's buildUsageOnlyChunk(id, model, usage)
  call shape; the fork's stale object form set `choices: []` as the usage payload.
diegosouzapw added a commit to initguru/OmniRoute that referenced this pull request Sep 15, 2026
…t gate behavior

Answers the open technical question from diegosouzapw#12902's review: does a
GLOBAL maxWaitMs=0 reintroduce the unbounded-queue regression diegosouzapw#12715
fixed (a request hanging ~6min until the client aborts)?

Evidence, exercising the real gate chatCore.ts actually calls
(accountSemaphore.acquireMany({ timeoutMs: requestQueue.maxWaitMs }),
not the Bottleneck reservoir the PR's own tests cover) under real
contention (maxConcurrency=1, two concurrent acquires):

  - No: it does not hang. setTimeout(reject, 0) fires on the next
    tick, so a second contending request is rejected with
    SEMAPHORE_TIMEOUT in low milliseconds, never minutes.
  - But it is also not a genuine 'no cap' — an operator setting 0
    expecting 'wait as long as it takes' instead gets near-zero
    tolerance for even momentary contention on any configured
    concurrency gate (global/provider/account). This is a real
    asymmetry vs. the Bottleneck reservoir path (where 0 truly means
    unbounded) left for the maintainer to decide how to resolve —
    not something this pass can decide unilaterally.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
diegosouzapw added a commit that referenced this pull request Sep 17, 2026
…expiration (#12902)

* fix(resilience): allow maxWaitMs=0 as disable sentinel for execution expiration

maxWaitMs normalization clamped the value to min:1, silently rewriting
an operator's 0 ("disable the limiter-managed execution deadline") into
1 — a 1ms expiration that killed every long-running job instantly. This
broke long-running reasoning models (GLM-5.2 with reasoning.effort=max
spends minutes before the first token, exceeding any practical
maxWaitMs; the TTB safety net is FETCH_TIMEOUT_MS, default 600s).

Fix: lower the floor to min:0 so 0 is preserved as the disable sentinel.
Issue #4165 follow-up.

Tests: 7/7 (resilience-normalize-maxwaitms-disable 5 + rate-limit-
maxwaitms-disable-execution 2). typecheck:core clean.

* fix(resilience): relax requestQueueSettingsSchema.maxWaitMs to allow 0

normalizeRequestQueueSettings already treats maxWaitMs=0 as an explicit
disable sentinel (queue-wait budget off), but the settings API schema
still rejected 0 with min(1), so an operator could never actually reach
the fix through PATCH /api/resilience. executionMaxWaitMs is untouched
(stays min(1) — separate field, separate decision, see #12902 item 4).

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* test(resilience): prove maxWaitMs=0 vs #12715's queue-wait gate behavior

Answers the open technical question from #12902's review: does a
GLOBAL maxWaitMs=0 reintroduce the unbounded-queue regression #12715
fixed (a request hanging ~6min until the client aborts)?

Evidence, exercising the real gate chatCore.ts actually calls
(accountSemaphore.acquireMany({ timeoutMs: requestQueue.maxWaitMs }),
not the Bottleneck reservoir the PR's own tests cover) under real
contention (maxConcurrency=1, two concurrent acquires):

  - No: it does not hang. setTimeout(reject, 0) fires on the next
    tick, so a second contending request is rejected with
    SEMAPHORE_TIMEOUT in low milliseconds, never minutes.
  - But it is also not a genuine 'no cap' — an operator setting 0
    expecting 'wait as long as it takes' instead gets near-zero
    tolerance for even momentary contention on any configured
    concurrency gate (global/provider/account). This is a real
    asymmetry vs. the Bottleneck reservoir path (where 0 truly means
    unbounded) left for the maintainer to decide how to resolve —
    not something this pass can decide unilaterally.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: Jihyun Son <jihyun.son@sk.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
diegosouzapw added a commit that referenced this pull request Sep 19, 2026
…ests + pack-policy + dashboard-typecheck)

Three production defects the tests caught:
- rateLimitManager: maxWaitMs=0 (the #12902 disable sentinel) hit #12715's
  queue-budget gate as "0 ms left" and 503'd every protected request.
- emergencyFallback: #14006 silently switched the budget-exhaustion target
  provider nvidia -> groq against ENVIRONMENT.md and the NIM snapshot; restored.
- claudeConnectionFields.ts vs ClaudeConnectionFields.tsx (#13074) differed only
  by casing; helpers renamed to claudeConnectionFieldValues.ts.

Guards realigned to legitimate changes: #13874 rotation map (distinct token in
the error test), #13350 origin-IP denylist, #13318 shared-catalog growth
(counts by invariant), comboTargetKeyPolicy import in the telegram stub, the
22 README mirrors that #13940/#14106 stamped with the retired openference.svg
(translated Cerebras cells recovered from history, hashes re-stamped),
bin/antigravity-bridge.mjs allowed in the pack policy, and the two dashboard
typecheck regressions (typed pinned section, ComponentProps cast).

Refs #13866.
diegosouzapw added a commit that referenced this pull request Sep 21, 2026
#12902 released requestQueue.maxWaitMs=0 as the sentinel that disables
the queue-wait deadline, but the #12715 queue-budget gate in
withRateLimit() (`if (queueRemainingMs <= 0) throw`) read 0 as "budget
spent" and rejected every request on a protected connection with an
immediate 503 queue-budget error — the exact opposite of what the
setting promises. rate-limit-maxwaitms-disable-execution ("400ms job
completes without 504") was red on the tip.

When no caller budget is passed and the configured queue budget is 0,
skip the gate, never arm the queue-wait timer and hand
awaitProviderDefaultSlot no budget (it falls back to the window).
Execution stays bounded by executionMaxWaitMs and the upstream
fetch-start timeout, as before.

Refs #13866
diegosouzapw added a commit that referenced this pull request Sep 21, 2026
Drains base-red waves 5–8 of release/v3.8.51: 35+ tests and the pack-policy,
api-typecheck, dashboard-typecheck, docs-all, agent-skills-sync, ESLint and
mutation-coverage gates (#13866).

Production defects the tests caught:
- Caveman: #12825's file-pack prefilter tested anchored rules against the
  original text, so leader_phrases never ran.
- pack-artifact: httpClientAbortGuard.mjs was missing from the staging
  allowlist and the required set — every published boot died with
  ERR_MODULE_NOT_FOUND (#14191).
- rateLimitManager: maxWaitMs=0 (the #12902 disable sentinel) hit #12715's
  queue-budget gate as "0 ms left" and 503'd every protected request.
- emergencyFallback: #14006 silently switched nvidia -> groq; restored per
  ENVIRONMENT.md and the NIM snapshot.
- ClaudeConnectionFields.tsx vs claudeConnectionFields.ts (#13074) differed
  only by casing; helpers renamed to claudeConnectionFieldValues.ts.
- /v1/responses/input_tokens body validated with Zod (HR#7).

Guards realigned to legitimate changes (#12663, #13863, #12565, #13990,
#13874, #13350, #13318, #13848 typed and split out of a size-capped file —
#14254), i18n catalogs for the 7 keys of #13074/#7f1b4a5e in 65 locales,
22 README mirrors restored to the Cerebras cell, env docs for 5 vars,
regenerated omni-version-manager skill.

Refs #13866. Closes #14254. Refs #14191.

Co-authored-by: Prabhjot Singh <jotgill1522@gmail.com>
Co-authored-by: Xmon Dai <xiechimon@qq.com>
diegosouzapw added a commit that referenced this pull request Sep 22, 2026
…ll, TS2677, ESLint (Refs #13866) (#14331)

* fix(build): ship httpClientAbortGuard.mjs in the pack artifact; validate input_tokens with Zod

Wave five of the release/v3.8.51 base-reds, part 1 — the two that matter.

#14064 restored server-ws.mjs's import of ./httpClientAbortGuard.mjs and the
assembleStandalone copy, but not the two pack-artifact policy entries that
were lost with it. Without APP_STAGING_ALLOWED_EXACT_PATHS the prepublish
prune deletes the file; without PACK_ARTIFACT_REQUIRED_PATHS nothing notices.
Every boot of the published package would die with ERR_MODULE_NOT_FOUND — the
3.8.47 head-response-guard class. Both closure suites (9/9) now enforce it.

#13910's /v1/responses/input_tokens read request.json() behind a hand-rolled
typeof check. Hard Rule #7 wants the boundary on Zod; the t06 guard caught it.
Same passthrough envelope the catch-all Responses route uses, since the
counters below already walk the fields defensively. 9/9 on the route's suite.

Five no-unused-vars left behind by the wave (cliRuntime execFileSync, arena
test symbols and a type, compression rmSync, waitForServer req) are removed.
The 'openwa routes removed without deprecation' entry from the #14101 run was
an artifact of that PR trailing its base — the gate is clean on the tip.

Refs #13866

* fix(compression): let anchored file-pack rules see the transformed text; align wave-5 guards

Wave five of the release/v3.8.51 base-reds, part 2.

One production defect. #12825 (Hungarian Caveman pack) stopped gating file-pack
rules with the English keyword list and tested the rule's own regex instead —
against `lowerResult`, a lower-cased copy of the ORIGINAL text that the loop
never refreshed. An anchored pattern like leader_phrases' `^(?:i will|…)`
therefore ran its prefilter on "sure, i will…", failed the anchor, and was
skipped; the rule that strips "I will " from every English response was dead
since the merge. The prefilter now sees the text as the rules so far have left
it. New test fails on the tip and passes here; all Caveman suites, Hungarian
included, are 95/95. A frozen no-unused-vars suppression on caveman.ts no
longer had a target and is pruned.

Two more TS2677 predicates of the kind #14101 fixed: #13910
(rerankProviderNodes.ts, `n is RerankProviderNodeRow` on a Record row) and
#13957's mitm catalog (antigravity.ts, `c is DynamicCatalogModel` on a
literal-or-null). Both narrow by NonNullable of the element's own type; the
api-route typecheck was 285 against a baseline of 283 on the pristine tip.

The rest are guards trailing legitimate changes:

- #12663 made gemini-3.8-flash the catalog head; T28 pinned 3.7.
- #13863 put mimo-v2.5 into the shared vision heuristic on purpose (the base
  model is multimodal, only the Pro variants are text-only). The safety test
  now asserts the real invariant: base and :free aliases yes, -pro no.
- #12565 moved npm-prefix detection into cliRuntimeNpmPrefix.ts with a
  process-lifetime cache that importFresh() does not reset; the case resets it.
  #12565 also builds Windows candidates with path.win32 on purpose; the qodercli
  test compared against POSIX path.join.
- #13990 (the 2 GB Docker image) copies better-sqlite3 with --chown; the guard
  matched the flag order literally. Now flag-order tolerant, still fails when
  --from=builder is removed.
- #13378 reintroduced public/openference.svg under a name #11750 retired for
  missing provenance and swapped the Cerebras showcase cell for it. The cell is
  back and the asset is gone; whether the new drawing counts as provenance is
  the owner's call.

Refs #13866

* fix(i18n): translate the 7 sidebar-pin and Claude low-priority keys into all 65 locales

#7f1b4a5e (sidebar pinned items) and #1b2349de (Claude OAuth lower-priority /
auto-reset) landed with their 7 new keys in en.json only, which the vi and
pt-BR parity suites flag. Translated with the repo's own sync-ui-keys
--translate-markers against the .113 i18n instance (codex/gpt-5.6-sol-low):
+446 lines across 65 catalogs, zero __MISSING__ markers, placeholders intact.
vi.json also has two keys reordered to mirror en.json; values unchanged.

Refs #13866

* chore(quality): list the 8 covering tests the sixth wave added in stryker tap.testFiles

30 commits landed on release/v3.8.51 while wave five drained; eight new unit
tests cover mutated modules and were not in tap.testFiles, so their mutant
kills did not count and check:mutation-test-coverage --strict failed on the
merged tree. Appended at the end of the list, nothing reordered.

Refs #13866

* docs: document the five env vars of the 09-18 wave; regenerate the version-manager skill for the open-wa routes

check:docs-all: BRIDGE_PORT, ROUTER_URL and CERT_DIR (bin/antigravity-bridge.mjs,
#c74cea3d), OPENWA_SERVICE_PORT (src/lib/services/bootstrap.ts, #1e8c913c) and
NEXT_PUBLIC_PORT (src/shared/hooks/useDisplayBaseUrl.ts, #d715190b) were read
in code but absent from .env.example and docs/reference/ENVIRONMENT.md. Added
next to their neighbours, with the defaults the code actually uses (open-wa is
8323, not the 201xx range the other services sit in).

check:agent-skills-sync: the open-wa feature added eight /api/services/openwa/*
routes to docs/openapi.yaml without regenerating skills/omni-version-manager/
SKILL.md. Regenerated with the repo generator; the diff is exactly those eight
route sections.

Refs #13866

* test: register the crash guard in the pack snapshot; inventory #13874's refresh-lane row read

pack-artifact-policy pins the list of root runtime files check:pack-artifact
must find in the tarball; dist/httpClientAbortGuard.mjs joined
PACK_ARTIFACT_REQUIRED_PATHS in this PR and the snapshot follows.

#13874 re-reads the connection row inside the Claude refresh lane so a queued
health check does not POST a refresh token a Layer 2 refresh already rotated —
a state read, inventoried like the family-cooldown lookup (tokenHealthCheck.ts
2 -> 3).

Refs #13866

* chore(quality): list native-codex-auto-resume test in stryker tap.testFiles (#13180 landed without it)

* fix(release): drain the seventh base-red wave of release/v3.8.51 (9 tests + pack-policy + dashboard-typecheck)

Three production defects the tests caught:
- rateLimitManager: maxWaitMs=0 (the #12902 disable sentinel) hit #12715's
  queue-budget gate as "0 ms left" and 503'd every protected request.
- emergencyFallback: #14006 silently switched the budget-exhaustion target
  provider nvidia -> groq against ENVIRONMENT.md and the NIM snapshot; restored.
- claudeConnectionFields.ts vs ClaudeConnectionFields.tsx (#13074) differed only
  by casing; helpers renamed to claudeConnectionFieldValues.ts.

Guards realigned to legitimate changes: #13874 rotation map (distinct token in
the error test), #13350 origin-IP denylist, #13318 shared-catalog growth
(counts by invariant), comboTargetKeyPolicy import in the telegram stub, the
22 README mirrors that #13940/#14106 stamped with the retired openference.svg
(translated Cerebras cells recovered from history, hashes re-stamped),
bin/antigravity-bridge.mjs allowed in the pack policy, and the two dashboard
typecheck regressions (typed pinned section, ComponentProps cast).

Refs #13866.

* test: type the #13848 Gemini pairing tests (no-explicit-any) and inventory the semantic-cache embedding picker's connection read

Both arrived with the tip merge: #13848 added 13 explicit any casts to
translator-openai-to-gemini.test.ts (no-explicit-any is an error under
tests/), and 7a92129's embeddingOptions.ts reads provider connections
once without a hard-session-lease inventory entry. Stale suppression
count pruned for the test file only.

Refs #13866.

* test: split the #13848 turn-pairing cases out of translator-openai-to-gemini.test.ts

The file sits exactly at its frozen size cap; typing the pairing tests
(no-explicit-any) pushed it 14 lines over. The two cases are a coherent
regression suite of their own, so they move to
translator-openai-to-gemini-turn-pairing-13848.test.ts (registered in
stryker tap.testFiles) instead of widening the baseline.

* docs(env): document BRIDGE_PORT, ROUTER_URL, CERT_DIR, OPENWA_SERVICE_PORT and NEXT_PUBLIC_PORT (Refs #13866)

check:env-doc-sync has been red on the release tip since these five vars
reached code without their .env.example / ENVIRONMENT.md entries:
bin/antigravity-bridge.mjs (BRIDGE_PORT, ROUTER_URL, CERT_DIR — #14006),
src/lib/services/bootstrap.ts + api/services/openwa/_lib.ts (OPENWA_SERVICE_PORT)
and src/shared/hooks/useDisplayBaseUrl.ts (NEXT_PUBLIC_PORT — #13533).
Defaults and source files copied from the reads themselves.

* chore(skills): regenerate omni-version-manager for the open-wa service routes (Refs #13866)

check:agent-skills-sync (Merge integrity job) has been red on the tip since
the open-wa embedded-service routes reached docs/openapi.yaml without the
generated SKILL.md being refreshed. Output of
scripts/skills/generate-agent-skills.mjs --apply, no hand edits: the eight
/api/services/openwa/* operations.

* fix(types): make the two TS2677 type predicates sound (Refs #13866)

check:api-typecheck has been red on the tip with two "type predicate's
type must be assignable to its parameter's type" errors:

- src/app/api/v1/_shared/rerankProviderNodes.ts (#13733): the read cache
  hands back `Record<string, unknown> | null`, and an interface whose members
  are all optional is not assignable to an index-signature type. Narrow to the
  non-null record and assert the row shape afterwards.
- src/mitm/handlers/antigravity.ts (#14006): the map callback returned
  `{ displayName: string }` while DynamicCatalogModel declares it optional, so
  the predicate could not be proven. Type the callback's return explicitly and
  filter on `!== null`.

No runtime change; rerank-remote-provider-nodes / rerank-local-node-shapes /
mitm-handler-antigravity stay green.

* fix(lint): clear the 92 ESLint errors the lint gate reports on the tip (Refs #13866)

- tests/unit/translator-openai-to-gemini.test.ts: #13848 / #13318 added 13
  `any` casts/params on top of the 74 frozen for the file, so ESLint reported
  all 87. Typed them (GeminiRequestWithContents / GeminiToolPart, and the
  existing GeminiRequestWithConfig) and pruned the file's suppression to the
  new count of 71 — nothing else in eslint-suppressions.json changes.
- no-unused-vars: execFileSync import (src/shared/services/cliRuntime.ts,
  #12565), getArenaEloSyncStatus + makeLeaderboardMap + ArenaLeaderboardMap
  (tests/unit/arena-elo-sync-redesign.test.ts, #13446), rmSync
  (compressionAnalyticsWriterFlatRate.test.ts, #13446), `req` → `_req`
  (waitForServer-slow-first-response.test.mjs).

translator-openai-to-gemini 48/48; arena-elo-sync-redesign,
compressionAnalyticsWriterFlatRate, waitForServer-slow-first-response green.

* fix(compression): stop skipping anchored Caveman rules that only match after earlier rules

#12825 (Hungarian pack) replaced the English keyword prefilter with a
`rule.pattern.test(lowerText)` pre-check for every file-based rule, including
the default `en` pack. `lowerText` is the ORIGINAL message, so anchored rules
such as `leader_phrases` (`^i will …`) — which only match after `pleasantries`
strips "Sure, " — were dropped before they could run. `caveman-v379` caught the
regression ("I will ensure …" survived at full intensity).

Tag file-based rules with their pack language in ruleLoader and let the
keyword prefilter apply to `en`/built-in rules only; non-English packs (which
reuse English rule names) simply run their localized regex, which is what the
pre-test cost anyway. Drops the now-unused CAVEMAN_RULES import and prunes the
already-stale `caveman.ts` no-unused-vars suppression (0 violations on the tip)
that blocked the pre-commit hook for any change to this file.

Refs #13866

* test(models): align catalog and vision-heuristic guards with the tip's intended contracts

Three base-reds where the production change was deliberate and the pinned
guard was simply not bumped by the PR that changed the contract:

- agy-antigravity-shared-catalog-12724: #13318 added the three Gemini 3.8
  Flash tiers (high/medium/low, no "-tiered" endpoint for 3.8) to the shared
  Antigravity/AGY base, 10 -> 13. Pin the new size in one constant and make the
  buildSurfaceCatalog delta assertions relative to it.
- t28-model-catalog-updates: #12663 (issue #12638) registered gemini-3.8-flash
  at the head of the AI Studio fallback catalog as the current Flash default;
  assert 3.8 first and keep 3.7 present.
- command-code-mimo-v2-5-safety: #13863 (issue #13847) added an explicit
  "mimo-v2.5" fragment to the shared vision heuristic so provider-qualified and
  `-free` aliases keep their vision flag. The guard's real concern (the
  "mimo-vl" fragment must not cover "mimo-v2.5") is asserted on the fragment
  itself; the bare id is now vision by heuristic on purpose, and the Pro
  text-only sibling stays excluded.

Refs #13866

* test(cli): follow the #12565 cliRuntime module split in the npm-prefix and qodercli guards

#12565 (issue #12563) moved the npm global-prefix cache out of cliRuntime.ts
into cliRuntimeNpmPrefix.ts and built the Windows known-bin candidates with
`path.win32` (cliRuntimeWindowsNode.ts) so they stay Windows-shaped when
`process.platform` is mocked on a POSIX runner. Two pre-existing guards
depended on the old layout:

- cli-runtime-extended "resolves known binaries from npm global prefix":
  importFresh() only re-evaluates cliRuntime.ts; the prefix cache now lives in
  a module that stays shared across cases, so a real `npm config get prefix`
  from an earlier case was cached and the mocked execFileSync never ran. Reset
  the cache with the helper #12565 exported for exactly this in afterEach.
- qodercli-windows-resolve-6263: compare against `path.win32.join` — identical
  to `path.join` on a real Windows host, which is the behaviour under test.

Production behaviour is unchanged on both platforms.

Refs #13866

* test(auto-update): write the source-mode log inside the test's own temp dir

The launchAutoUpdate case pointed AUTO_UPDATE_LOG_PATH at a fixed, world-shared
`/tmp/auto-update-source.log`. On the .113 runner the suite executes both as
`root` and as `runner` (uid 1001): the file survives owned by whoever ran
first (`-rw-r--r-- root root`), and the next `openSync(logPath, "a")` fails
with EACCES for the other user. Reproduced locally by making the shared file
read-only; production code is untouched (autoUpdate.ts last changed in #9354).

Use a per-test mkdtemp path for the source-mode log and clean the whole temp
root in the existing finally block.

Refs #13866

* fix(dashboard): rename claudeConnectionFields.ts so it no longer case-collides with ClaudeConnectionFields.tsx

#13074 added two modules to the provider-detail modals directory whose names
differ only by casing: `ClaudeConnectionFields.tsx` (the component) and
`claudeConnectionFields.ts` (the value/patch helpers). On a case-insensitive
filesystem the pair breaks the webpack build (#6584 guard), and esbuild's
resolver already picks the `.tsx` for the extension-less `./claudeConnectionFields`
specifier, so the provider-detail client entry failed to bundle ("No matching
export ... for import claudeConnectionFieldPatch").

Rename the helper module to `claudeConnectionFieldValues.ts` (the same naming
the sibling `quotaScrapingFieldValues.ts` uses) and point the only importer,
EditConnectionModal.tsx, at the new name. Greens
tests/unit/case-collision-6584.test.ts and
tests/unit/media-page-client-browser-bundle.test.ts.

Refs #13866

* fix(build): allowlist dist/httpClientAbortGuard.mjs so the published tarball keeps the server-ws crash guard

#14064 (re-land of #13636) made scripts/dev/standalone-server-ws.mjs import
./httpClientAbortGuard.mjs and taught assembleStandalone to copy the shared
implementation next to dist/server-ws.mjs — but never registered the file in
scripts/build/pack-artifact-policy.ts. The prepublish prune deletes anything
outside APP_STAGING_ALLOWED_EXACT_PATHS, and check:pack-artifact only fails on
PACK_ARTIFACT_REQUIRED_PATHS entries, so the next `omniroute` tarball would
boot straight into ERR_MODULE_NOT_FOUND (the #7065 / tls-options class the
closure tests exist to catch).

Add the bare and dist/ entries to both lists and extend the required-paths
snapshot in tests/unit/pack-artifact-policy.test.ts. Greens
tests/unit/pack-artifact-entrypoint-closures.test.ts and
tests/unit/pack-artifact-server-ws-closure.test.ts.

Refs #13866

* test(docker): accept --chown=node:node on the better-sqlite3 runner COPY

#14010 deliberately changed the runner-stage COPYs to `COPY --chown=node:node
--from=builder ...` (ownership at copy time instead of a second ~2 GB
`chown -R` overlay layer). The Dockerfile contract test still matched the old
`COPY --from=builder /app/node_modules/better-sqlite3` prefix and went red on
the tip even though the native-addon guard it protects is intact. Tolerate the
optional --chown flag; every other assertion (node-gyp rebuild, both
`test -f .../better_sqlite3.node` checks) is unchanged.

Refs #13866

* fix(api): validate /v1/responses/input_tokens bodies with Zod (t06)

#13167 added the local Responses token-count route with hand-rolled
`typeof` checks on `request.json()`. Hard Rule #7 and the t06 gate
(scripts/check/check-route-validation.mjs, mirrored by
tests/unit/route-body-validation-t06.test.ts) require every route that reads
request.json() to go through validateBody()/safeParse(), so the tip was red.

Add `v1ResponsesInputTokensSchema` (pins the wire types the counter reads —
model/instructions strings, input string-or-array, tools array — and lets
unknown keys through since they are counted, never forwarded) and run the body
through validateBody(); a type mismatch is now a 400 naming the field instead
of a silently ignored key. Regression test added to
tests/unit/responses-input-tokens-local-route.test.ts.

Refs #13866

* fix(docs): drop the retired openference.svg asset reintroduced by #13378

`openference.svg` is one of the 78 provider assets retired for missing
provenance (tests/unit/provider-assets-generic-fallback.test.mjs freezes that
list and forbids any tracked surface from referencing a retired name). #13378
added a new hand-drawn `public/openference.svg` outside the manifest-audited
public/providers/ tree and pointed the README free-tier table (plus the 22
i18n mirrors that carry the row) at it, which put the retired name back on a
tracked surface and left an unaudited asset in the package.

Use the generic fallback icon (`public/providers/cli-generic.svg`) the other
provenance-less providers already use, delete the unaudited file, and adopt
the mechanical README edit into .i18n-state.json
(`i18n:run -- --adopt --files=README.md`, no API calls) so the i18n drift gate
does not flag README.md as source-changed.

Refs #13866

* test(lease): classify the two connection-query sites added by #14159 and #13874

The hard-lease bypass inventory froze every getProviderConnections /
getProviderConnectionById site with a class; two landed on the tip without a
golden update:

- src/app/api/settings/cache-config/embeddingOptions.ts (#14159, re-land of
  #12630): read-only listing that feeds the semantic-cache embedding dropdown,
  same shape as the qdrant embedding-models route — class C.
- src/lib/tokenHealthCheck.ts 2 -> 3 (#13874): re-reads the row by id after an
  unrecoverable refresh error to detect credentials rotated by a concurrent
  Layer 2 refresh before deactivating — a state read, not dispatch; stays C.

Refs #13866

* fix(sse): restore nvidia as the emergency budget-fallback provider

#14006 (Antigravity MITM catalog injection) flipped
EMERGENCY_FALLBACK_CONFIG.provider from "nvidia" to "groq" in one line,
without touching ENVIRONMENT.md, .env.example, the chat.ts comment or the
NVIDIA hosted-model snapshot, all of which still promise
nvidia/openai/gpt-oss-120b. Operators without a Groq connection got the
original 402 back instead of the free reroute, and
chat-route-coverage ("uses the emergency fallback model on budget
exhaustion" / "returns the primary budget error when emergency fallback
also fails") went red on the tip.

Put the documented default back; the #14006 bridge tests exercise
bin/antigravity-bridge.mjs and do not read this config.

Refs #13866

* fix(resilience): keep maxWaitMs=0 a "no queue deadline" sentinel

#12902 released requestQueue.maxWaitMs=0 as the sentinel that disables
the queue-wait deadline, but the #12715 queue-budget gate in
withRateLimit() (`if (queueRemainingMs <= 0) throw`) read 0 as "budget
spent" and rejected every request on a protected connection with an
immediate 503 queue-budget error — the exact opposite of what the
setting promises. rate-limit-maxwaitms-disable-execution ("400ms job
completes without 504") was red on the tip.

When no caller budget is passed and the configured queue budget is 0,
skip the gate, never arm the queue-wait timer and hand
awaitProviderDefaultSlot no budget (it falls back to the window).
Execution stays bounded by executionMaxWaitMs and the upstream
fetch-start timeout, as before.

Refs #13866

* test: align three fixtures with the #13874, #13861 and #13350 contracts

Three base-reds that are deliberate contract changes, not defects:

- executor-default-base "refreshCredentials swallows refresh errors":
  #13874 records rotations on the Layer 2 (no connectionId) refresh path
  too, so the "refresh-me" token the previous case already rotated was
  served from the rotation map without the network POST the test wanted
  to fail. Use a token nobody rotated.
- telegram-keycache-bounded-13165: #13861 made comboTargetKeyPolicy
  import isModelBlockedByPatterns from db/apiKeys; the loader-stubbed
  module lacked it and the suite died at module load. Export an honest
  "not blocked" stub (the test has no blocked models).
- upstream-headers-proxy-auth "ordinary headers are still allowed":
  #13350 forbids the whole origin-IP forwarding set upstream (covered
  by upstream-headers-sanitize). Swap x-forwarded-for for x-request-id.

Refs #13866
@maxmad64bis
maxmad64bis deleted the feat/cascade-pr1-resilience branch September 24, 2026 21:11
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…iegosouzapw#12715)

Fila sem teto que segura a request seis minutos até o cliente abortar é pior que 503 imediato: consome slot, mascara a saturação e ainda entrega erro no fim. Um orçamento `maxWaitMs` por conexão compartilhado entre gate, slot padrão do provider e fila do Bottleneck é a forma certa — o teto tem que ser um só, senão cada camada espera o seu.

O `max(perConn, upstream)` no `executionMaxWaitMs` é o detalhe que evita a correção matar request em voo, que seria trocar um defeito por outro.

Registro a atribuição: você manteve o diegosouzapw#12635 aberto para o @Tushar49 e creditou a percepção dele (providers lentos precisam de 2min→10min por conexão) enquanto adiciona o encanamento que faltava. É o jeito certo de construir sobre PR de outra pessoa sem tomar o crédito.

Sobre o `npm run lint` desmarcado com a nota do eslint quebrado no ambiente: deixar em branco e explicar vale mais que marcar sem ter rodado. Rodei aqui: limpo.

Revalidei sobre o tip: **13/13**, typecheck:core limpo, check-file-size OK. O `file-size-baseline.json` conflitou com os rebaselines desta campanha — resolvido aditivamente, JSON revalidado com `json.load`.
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…expiration (diegosouzapw#12902)

* fix(resilience): allow maxWaitMs=0 as disable sentinel for execution expiration

maxWaitMs normalization clamped the value to min:1, silently rewriting
an operator's 0 ("disable the limiter-managed execution deadline") into
1 — a 1ms expiration that killed every long-running job instantly. This
broke long-running reasoning models (GLM-5.2 with reasoning.effort=max
spends minutes before the first token, exceeding any practical
maxWaitMs; the TTB safety net is FETCH_TIMEOUT_MS, default 600s).

Fix: lower the floor to min:0 so 0 is preserved as the disable sentinel.
Issue diegosouzapw#4165 follow-up.

Tests: 7/7 (resilience-normalize-maxwaitms-disable 5 + rate-limit-
maxwaitms-disable-execution 2). typecheck:core clean.

* fix(resilience): relax requestQueueSettingsSchema.maxWaitMs to allow 0

normalizeRequestQueueSettings already treats maxWaitMs=0 as an explicit
disable sentinel (queue-wait budget off), but the settings API schema
still rejected 0 with min(1), so an operator could never actually reach
the fix through PATCH /api/resilience. executionMaxWaitMs is untouched
(stays min(1) — separate field, separate decision, see diegosouzapw#12902 item 4).

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* test(resilience): prove maxWaitMs=0 vs diegosouzapw#12715's queue-wait gate behavior

Answers the open technical question from diegosouzapw#12902's review: does a
GLOBAL maxWaitMs=0 reintroduce the unbounded-queue regression diegosouzapw#12715
fixed (a request hanging ~6min until the client aborts)?

Evidence, exercising the real gate chatCore.ts actually calls
(accountSemaphore.acquireMany({ timeoutMs: requestQueue.maxWaitMs }),
not the Bottleneck reservoir the PR's own tests cover) under real
contention (maxConcurrency=1, two concurrent acquires):

  - No: it does not hang. setTimeout(reject, 0) fires on the next
    tick, so a second contending request is rejected with
    SEMAPHORE_TIMEOUT in low milliseconds, never minutes.
  - But it is also not a genuine 'no cap' — an operator setting 0
    expecting 'wait as long as it takes' instead gets near-zero
    tolerance for even momentary contention on any configured
    concurrency gate (global/provider/account). This is a real
    asymmetry vs. the Bottleneck reservoir path (where 0 truly means
    unbounded) left for the maintainer to decide how to resolve —
    not something this pass can decide unilaterally.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: Jihyun Son <jihyun.son@sk.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants