fix(resilience): stop unbounded queue that hangs 6min until Aborted - #12715
Merged
diegosouzapw merged 27 commits intoSep 10, 2026
Merged
diegosouzapw merged 27 commits into
diegosouzapw merged 27 commits into
Conversation
maxmad64bis
force-pushed
the
feat/cascade-pr1-resilience
branch
from
September 4, 2026 08:12
e44c35d to
f9c7e60
Compare
maxmad64bis
marked this pull request as ready for review
September 4, 2026 08:50
maxmad64bis
force-pushed
the
feat/cascade-pr1-resilience
branch
2 times, most recently
from
September 4, 2026 08:55
ac18315 to
250c39a
Compare
Per-connection maxWaitMs budget is now shared across acquisition gates, provider-default slot wait, and Bottleneck queue (fail-closed 503 on exhaust). Prevents the 45s gate+slot+queue sum that previously hung requests until manual abort. Execution backstop executionMaxWaitMs is now overridable per connection and clamped to the real upstream fetch timeout so it never kills a mid-flight request.
maxmad64bis
force-pushed
the
feat/cascade-pr1-resilience
branch
from
September 10, 2026 11:51
250c39a to
2061fd2
Compare
Contagem de migrations 171 → 172 após a `175_call_logs_provider_stats_indexes.sql` do diegosouzapw#12832. Medido com `ls src/lib/db/migrations/*.sql | wc -l`. Falha minha de processo: depois da onda 1 desta leva eu medi file-size, api-typecheck, changelog-integrity e colisão de migration — não o `check:docs-counts`. O drift ficou vivo até a onda 2 esbarrar nele. 41 mirrors de `llm.txt` regenerados pelo script do projeto. Aprovado por você para tocar `AGENTS.md`, mesma classe do diegosouzapw#12970.
…thout a documented free tier (diegosouzapw#12744) Validado numa worktree combinada com a onda de roteamento/free-tier desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (288), check-file-size OK após rebaseline, 84/85 nos testes focados no runner Node. Parar de honrar `isFree:true`, `:free` e `0/0` vindos do upstream **antes** de consultar o catálogo é a inversão certa: hoje um provider sem tier livre documentado consegue se declarar grátis e o listing diverge do roteador `auto/*`. Checar as heurísticas depois do hit de catálogo fecha a porta sem quebrar o caminho de linhas custom locais, que continuam confiáveis pelo caminho próprio. O `isFreeModel("or", …)` com um alias que não existe é o tipo de bug que passa despercebido porque falha silenciosamente para o lado permissivo. **Integração:** o `decideHidePaid` que o diegosouzapw#12795 extraiu passou a usar o seu `isFreeForProvider` por id, em vez de OR-ear os dois aliases num único `freeProvider`. A forma por id é a garantia que esta PR estabelece, então ela prevaleceu.
…nifest (diegosouzapw#12786) Validado numa worktree combinada com a onda de roteamento/free-tier desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (288), check-file-size OK após rebaseline, 84/85 nos testes focados no runner Node. Metadado de display derivado da mesma fonte da decisão, com um gate STRICT que quebra o CI se a contagem do manifesto divergir do catálogo — é o detalhe que impede a tag de virar mentira daqui a três meses. Registrar que 77 entradas de catálogo viram 76 no manifesto porque o `arcee-ai` ainda não tem entrada no registry é exatamente o tipo de discrepância que costuma virar bug fantasma.
…ty/breaker (diegosouzapw#12794) Validado numa worktree combinada com a onda de roteamento/free-tier desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (288), check-file-size OK após rebaseline, 84/85 nos testes focados. Tratar breaker aberto como equivalente a fechado no scoring de snapshot é pior que não pontuar: afirma saúde onde há falha conhecida. Trocar as três constantes neutras por valor observado é a correção, e manter preço e orçamento fora do escopo mantém a PR revisável. **Integração:** o `computeSnapshotWeights` conflitou com o diegosouzapw#12731, que adiciona peso de `reliability` a partir de `failureRate`/`errorRate`. Os dois cobrem chaves diferentes e compõem — ficaram ambos: reliability do diegosouzapw#12731, health via breaker e quality desta PR, quota neutro nos dois. Nenhum dos dois lados foi descartado.
…diegosouzapw#12731) Validado numa worktree combinada com a onda de roteamento/free-tier desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (288), check-file-size OK após rebaseline, 84/85 nos testes focados. `reliability-first` que não pesava reliability é o defeito mais constrangedor possível num mode pack, e a causa é clara: `modePacks.ts:13` substituía os defaults por inteiro. Financiar os novos pesos com `quota`/`costInv`/`tierPriority` mantendo cada pack somando 1.0 é a parte que exige cuidado e você fez. Manter `quality-first` em 0.03, igual ao default, para que ele não fique mais fraco que `balanced`, é o tipo de detalhe que só aparece quando se checa a tabela inteira. **Integração:** conflitou com o diegosouzapw#12794 no `computeSnapshotWeights`; os dois compõem e ambos ficaram.
…iegosouzapw#12790) Validado numa worktree combinada com a onda de roteamento/free-tier desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (288), check-file-size OK após rebaseline, 84/85 nos testes focados. Duas cópias da mesma ordem de provider com um "keep in sync" implícito é dívida que cobra juros a cada provider novo. Uma definição com re-export nos dois lados resolve a classe. O `xao/*` ordenando depois de todo provider conhecido em vez de junto do `xai-oauth` é um sintoma concreto de que a duplicação já estava divergindo. **Integração:** `scripts/quality/run-all-gates.mjs` conflitou com o `check:pricing-freshness` que entrou pelo diegosouzapw#12792 na mesma onda. Aditivo — os dois gates coexistem.
…s, pooled latency bootstrap, fresh tier cache (diegosouzapw#12792) Um modelo grátis fora da tabela herdando $5/$15 por milhão e afundando no roteamento cost-aware é o defeito mais caro desta onda: silencioso, e inverte exatamente a decisão que o operador quer. Parar de chutar 1500ms de latência para modelo desconhecido e usar a mediana observada do pool — com contador de quantas vezes o chute dispara — é trocar heurística por medição do jeito certo. O contador é o que permite saber se valeu. Revalidei após reconstruir a branch sobre o tip: **33/33** nas suítes da PR, typecheck:core limpo, `check-api-typecheck` OK (289). **Duas integrações:** 1. `computeSnapshotWeights` conflitou com o diegosouzapw#12794 (health via breaker + quality), já mergeado. Os dois compõem e ambos ficaram: o seu termo de `reliability` — que era a única chave que o caminho de snapshot ainda ignorava — mais o health observado e o quality do diegosouzapw#12794. 2. `scripts/quality/run-all-gates.mjs` conflitou com o `check:provider-order-sync` do diegosouzapw#12790. Aditivo, os dois gates coexistem. **Nota de dívida:** o `virtualFactory.ts` cruzou o teto de 1200 linhas pela primeira vez (1187 → 1207) somando esta onda. Congelei em vez de dividir e registrei os dois candidatos a extração na justificativa — `computeSnapshotWeights` (~85 linhas) e o grupo de elegibilidade de credencial (~70). Qualquer um dos dois volta o arquivo para baixo do cap.
…op reasons (diegosouzapw#12795) Um pool `auto/*` vazio que não diz por que está vazio é a pior forma de falha: o operador vê ausência e não sabe se é config, cota ou catálogo. Registrar qual estágio removeu quantos, e carregar `dropReason` em cada entrada de `quotaHealth.providers`, transforma silêncio em diagnóstico. Os três defeitos achados de carona valem tanto quanto a feature — em especial a cota livre recorrente sem teto sendo tratada como desconhecida em vez de segura, que é justamente o caso que mais aparece. Revalidei após reconstruir sobre o tip: **15/15**, typecheck:core limpo, `check-api-typecheck` OK (289). **Uma mudança minha no `catalogPaidFilter.ts`.** Mantive a sua extração — ela é mais limpa que o predicado inline — mas troquei o corpo para usar `isFreeForProvider` por id. O módulo OR-eava `providerHasFreeModels(resolved) || providerHasFreeModels(canonical)` num único `freeProvider` e depois testava `isFreeModel` em cada um. Com isso, um id cujo próprio provider não documenta tier livre passa a ler como grátis sempre que o alias irmão documenta — que é exatamente o buraco que o diegosouzapw#12744 fechou e já está no tip. Por id preserva a garantia; nenhuma outra linha do módulo mudou. **Dívida registrada:** o `virtualFactory.ts` foi de 1187 para 1219 somando esta onda e cruzou o teto de 1200 pela primeira vez. Congelado com os candidatos a extração nomeados na justificativa.
…diegosouzapw#13044) Batch 1 of the locale expansion: Greek, Croatian, Serbian, Lithuanian, Estonian, Latvian, Slovenian, Maltese and Irish across the dashboard catalog, docs mirrors, CLI catalog, README, locale index and the site. 42 → 51 locales. Also fixes the ICU literal escape the translation backend dropped around angle placeholders, four translations that invented or renamed a placeholder, the language bars that linked to mirrors that do not exist, and the migration count drift (171 → 172).⚠️ base-red inherited: diegosouzapw#12732 — the four unit shards and Fast Quality Gates fail identically on unrelated PRs cut from the same base.
…e History tab (2.9) (diegosouzapw#12677) * feat(dashboard): pure model to compare two orchestration runs * feat(dashboard): compare-mode selection in the History grid * feat(dashboard): side-by-side comparison panel in the History tab (2.9) * chore(dashboard): compare-runs i18n + changelog * fix(dashboard): compare-panel loading state, height bound, delta legend, ARIA level Final-review fix wave for PR-A (Orchestration Canvas Fase 3): - Give each compare-panel side an explicit fetch status (loading/ok/error) so the Events metrics row shows "—" instead of a misleading real "0"/delta while a side is still loading or after its fetch failed. - Bound the compare panel's height (max-h-[45vh], overflow-y-auto, shrink-0) so a run with many activities can no longer collapse the History grid to zero height. - Clear the compare selection when the History preset changes, since a stale pick can fall outside the new range. - Add a delta-column legend (new compareDeltaLegend i18n key, translated into all 41 non-English locales) so operators know which side a positive delta favors. - Use role="status" (not role="alert") for the informational compareDifferentIdentity banner, reserving role="alert" for actual per-side fetch failures. - Restore ro.json's compareCost to the true cognate "Cost" (was distorted to "Cheltuieli" to dodge a byte-identical-to-English heuristic); audited the other 40 locales for the same pattern across the 9 compare* keys, no other instance found. * refactor(dashboard): split compare-runs/history-tab functions to clear complexity ratchets CompareRunsPanel (92 lines, max-lines-per-function) and HistoryTab (85 lines, same rule) exceeded the 80-line function cap; compareRuns.ts's a2aEventsFrom exceeded the cognitive-complexity cap (16 > 15). Extract pure/presentational helpers (ComparePanelHeaderBar, SideErrorRow, ComparisonMetrics, a2aEventFrom, HistoryStatusRows, refreshNowMsOnActionDone) with no behavior, DOM, i18n or aria change. --------- Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
…zapw#12392) (diegosouzapw#12983) * fix(dashboard): keep the first failure timestamp in sourceStale buildSourceStatuses stamped nowIso on every failing source at every poll, so the stale indicator reported "since the last poll" instead of the first failure — and, because snapshotContentKey serializes sources, the snapshot identity churned on every tick while any source was down. The failing branches now reuse the staleSince already held by that source in the previous status list, via the functional setStatuses updater (no ref read during render, no setState inside an effect body). Refs diegosouzapw#12392 * fix(dashboard): flag a source that starts failing after it had data buildRootAndSourceEdges only materialized a placeholder SourceNode when the failing source had no node at all. A source that already had work nodes and then started failing (or went offline) kept its healthy-looking SourceNode forever: no ⚠, no stale styling, no `sourceStale` line — the operator saw a normal source while it was actually broken. Now every non-ok/offline source is flagged: when its SourceNode is missing the placeholder is created as before; when it exists, the node is replaced by a copy carrying `sourceIssue` and `staleSince`. The copy (never a mutation) keeps the function pure — the original object is still referenced by the caller's `parts`, the same trap the droppedByState aliasing fix covered. Tests: three cases in tests/unit/ui/orchestrationModel.test.ts — existing node starting to fail (flags set, work nodes kept, no duplicate node, input object untouched), existing node going offline (no invented staleSince), and a healthy source staying free of both fields. Refs diegosouzapw#12392 * fix(dashboard): canvas polish batch (diegosouzapw#12392) Seven pointwise fixes on the Orchestration Canvas, each covered by a test: 1. Debounce x chip race: every chip/clear write in OrchestrationToolbar now cancels the pending search timer first. Left armed, it fired ~300ms later with a setParams closed over the pre-chip query string and silently reverted the chip. 2. The search input carries an aria-label (searchPlaceholder) — the placeholder alone is not an accessible name. 3. parseCsvSet trims each token, so `?state=running, failed` parses like the unpadded form instead of dropping the padded value. 4. toggleCsv was duplicated in the toolbar and the page client; both now import the single definition from the new model/urlParams.ts (pure, never mutates its inputs). 5. AgentsTab tells "nothing running" apart from "the filter matched nothing": with an active filter and no work node it renders noMatches + a clear-filters button instead of the setup CTAs, which would be wrong advice there. 6. Particle cap: orchestrationToFlow stamps `particles` on every edge and turns it off above PARTICLE_EDGE_CAP (40) simultaneously active edges — StatusEdge then renders the colored stroke without its 3 SMIL particles per edge. 7. The drawer's error banner clears when an action succeeds, so a recovered failure does not stay on screen. Only `noMatches` is added to en.json here; the other locales are task B4. Refs diegosouzapw#12392 * chore(dashboard): canvas polish i18n + changelog Real translations for orchestration.noMatches in the 41 non-English locales, each one written against that file's own neighbouring keys (emptyTitle, stateRunning, searchPlaceholder) so the wording for "task" and "filter" matches what the locale already uses. No i18n:sync-ui, no __MISSING__ left. Adds the changelog fragment for the nine PR-B fixes. Closes diegosouzapw#12392
…ck (diegosouzapw#12880) Validado numa worktree combinada com a onda de streaming desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 88/88 nos testes focados. `agent_message` chegando num fallback de Chat Completions é um item que o cliente não sabe interpretar; mapear ou descartar é a escolha certa, e escolher por item em vez de derrubar a resposta inteira mantém o fallback útil.
…ream stays silent (diegosouzapw#12828) Validado numa worktree combinada com a onda de streaming desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 88/88 nos testes focados. O diegosouzapw#12151 cobriu só metade: passthrough emitia o chunk final de usage, translate calculava a estimativa **depois** de fechar o stream, então o número só chegava ao log do servidor e nunca ao cliente. Fechar essa metade é o que faz a feature existir de fato. Não emitir segundo chunk quando o upstream já mandou usage real é o detalhe que impede a correção de virar contagem dobrada.
…p-dropped event array (diegosouzapw#12718) Validado numa worktree combinada com a onda de streaming desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 88/88 nos testes focados. Reconstruir o resumo a partir de um array que o próprio coletor já truncou por cap produz um resumo que parece completo e não é — pior que resumo ausente, porque não se distingue. Parar de reconstruir dali é a correção.
…ycle/heartbeat events, never real content (diegosouzapw#12741) Validado numa worktree combinada com a onda de streaming desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 88/88 nos testes focados. Um stream que só emite eventos de ciclo de vida e heartbeat, sem conteúdo nenhum, é falha disfarçada de sucesso: o cliente espera até o timeout dele. Falhar rápido devolve o controle.
…2941) Validado numa worktree combinada com a onda de streaming desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 88/88 nos testes focados. Estender a rotação que já existe para 429 ao 403 de bloqueio geográfico é a generalização certa, e manter a rejeição de fingerprint (Cloudflare 1010) fora dela é o que impede a rotação de queimar todas as contas contra uma recusa que não é de egresso. Nota: os checkboxes de validação do corpo ficaram em branco, mas o diff traz dois arquivos de teste — vale marcar da próxima para o revisor não precisar conferir.
…n in-memory pending store (diegosouzapw#12854) Diagnóstico por captura de pacote em tráfego real, com o `400 previous_response_not_found` reassemblado do tcpdump três vezes no mesmo loop de tool-calling — isso é evidência, não hipótese. A causa é limpa: `detail_state` só vira `'ready'` depois de uma escrita fire-and-forget enfileirada num worker único, e o cliente já tem o id de resposta antes disso. Semear a ponte **antes do primeiro await** é o que faz a correção não custar latência. Revalidei sobre o tip: **19/19**, typecheck:core limpo, check-file-size OK. **Estava draft e eu marquei como ready.** Não havia gate declarado — nem RFC pendente, nem decisão de produto em aberto — e passou na validação; a diretiva permanente do dono para esta campanha é avaliar draft como qualquer PR e promover quando passa. Se a intenção era segurar por outro motivo, me avise que eu reverto. **Um conserto meu na sua branch.** O `typecheck:core` falhava com `TS2345` em `callLogs.ts:489` — e falhava **na sua branch sozinha**, não por interação com a onda; confirmei isolando. O call site fazia cast para `{ clientRawRequest?: unknown; clientResponse?: unknown }`, mais frouxo que o `ContinuationPipeline` que o parâmetro exige, e `unknown` não assina para os membros tipados. Exportei o `ContinuationPipeline` do próprio store e usei ele no cast, em vez de alargar o tipo do parâmetro: o contrato passa a ter um nome só, no lugar onde ele já vivia. **Sobre a sua Reviewer Note do Map sem limite de contagem:** concordo que vale registrar. Entradas pequenas com TTL de 60s auto-expirando não justificam sizing agora, mas se aparecer burst sustentado o sintoma será memória, não erro — e aí a nota está aqui. Também carreguei o rebaseline de `chatCore.ts` (6021→6026) e `stream.ts` (3080→3098), que a onda de streaming inteira faz crescer.
…nal (diegosouzapw#12639) (diegosouzapw#12988) * feat(api): hydrate memoryHits from the persisted history event `GET /api/a2a/tasks/[id]` falls back to the persisted history row once a task leaves the in-memory TTL window, and `reconstituteHistoricalTask` hard-coded `metadata: {}` — so the drawer's "Memory used" section vanished for any historical task, even though `executeA2ATaskWithState` had already written a `memory_hits` event with the hits. The fallback now reads that event: `data_json` is parsed and, when it yields at least one well-formed hit, exposed as `metadata.memoryHits`. The event itself is filtered out of `events` — it is observability, not a state transition, and without the filter it leaked into the timeline as a duplicate of the row's current state. Reading is defensive throughout, mirroring `DrawerMemory`'s own validation: the payload is caller-influenced and unvalidated end to end, so `JSON.parse` runs inside `safeJsonParse`, non-arrays are rejected, and each entry must carry `id`, `key`, `type` and `snippet` as strings (a non-string field would be rendered as a React child and take the drawer down). Malformed input degrades to `metadata: {}` and a 200 — never a 500. Refs diegosouzapw#12639 * fix(a2a): bound the memory recall with its own deadline collectMemoryHits() runs BEFORE the skill handler and had no deadline at all, so a slow memory backend delayed the start of every A2A task — the HTTP genericBackend alone defaults to a 30s timeout. The search now races a MEMORY_RECALL_TIMEOUT_MS (1500ms) deadline. Overshooting degrades exactly like any other recall failure: empty hits, a warn log, and the task proceeds normally (best-effort contract unchanged, nothing propagates). The deadline timer is cleared in a finally on BOTH paths so no handle is left holding the event loop open, and MemoryHitsDeps.timeoutMs makes it injectable so the tests cost milliseconds instead of 1.5s of wall clock. Refs diegosouzapw#12639 * fix(dashboard): carry conductor requirements and focus the repeated task The drawer's "Repeat" for a Conductor task dropped the runner/model pinning and left the operator staring at the finished run: - `hubTaskSchema` now parses the hub's `requirements` (`.catch(null)` so an odd shape never fails the whole task parse), and `ConductorTaskDetail` exposes `cli`/`model` (`null` when the hub sends none). - `repeatReqForConductor` carries `cli`/`model` when present and OMITS them otherwise — the route's Zod takes both as optional strings, so a `null` would 400. The two fields are independent. - `performAction` reads the response body once and returns it, so the repeat can report the CANVAS id of the created task (`task_id` / `data.id` / `result.task.id`, each with its node prefix). `OrchestrationPageClient` then refetches and focuses it via `?node=`; History keeps its current behavior. - `conductor-routes-auth.test.ts` covers the creation route through its `ROUTES` array; the duplicated source assertion left `conductor-create-route.test.ts`. Refs diegosouzapw#12639 * chore(a2a): follow-ups changelog Changelog fragment for the five items PR-C delivers from diegosouzapw#12639. The sixth item on the issue — an authenticated panel path for A2A task creation — stays deliberately out of scope and is recorded as such in a comment on the issue rather than silently dropped: the JSON-RPC endpoint accepts API keys only, and widening that endpoint's auth surface to serve a UI convenience is the operator's call, not the implementation's. Closes diegosouzapw#12639
… across flow surfaces (diegosouzapw#12378) (diegosouzapw#13203) * refactor(ui): move shared flow colors to the orchestration status tokens FLOW_EDGE_COLORS and TokenHealthBadge were pinned to the fixed dark-mode hexes in STATUS_HEX, so both rendered dark-theme green/amber/red on a light background. They now read the theme-aware --orch-status-{success,warning, error,muted} custom properties introduced in Fase 2. The dark values of those tokens are exactly the old hexes, so dark mode is unchanged and only light mode gains contrast. `idle` was already a CSS var, which is the precedent proving a var() resolves in a ReactFlow edge stroke. Five call-sites built translucent variants by concatenating an 8-bit alpha suffix onto the palette hex (`${FLOW_EDGE_COLORS.error}40`), which cannot work with a var(). They move to a new documented helper, flowColorAlpha(), that wraps color-mix() — the same approach orchStateBadgeBg() already uses in the orchestration model. Percentages mirror the old suffixes (20 -> 13%, 30 -> 19%, 40 -> 25%). STATUS_HEX stays exported as the dark-mode mirror; it now has no production consumer. globals.css needed no change — all five tokens already existed in both themes. The colour assertions in the topology, combo-live and design-grid suites were aligned to the tokens, never weakened: every hex equality became an equality against the corresponding var(). design-grid additionally now asserts each token is defined in BOTH themes. Refs diegosouzapw#12378 * refactor(ui): finish the status-token migration across flow surfaces Sweeps the five state hexes across the remaining flow surfaces, following D1: - ComboLiveStudio: active/error provider pills and the run-outcome tri-state. - CompressionCockpit / WaterfallInspector / IoNode: the savings readouts and the savings quality ramp (>=30 success, >=15 warning, else muted). - EngineNode: the same ramp, plus the running state, whose glow moved to flowColorAlpha — the literal #f59e0b40 suffix is invalid once the value is a var(). - WaterfallInspector: a skipped step now reads as muted rather than a bare grey hex. Deliberately NOT migrated, because they are categorical or brand palettes rather than state: STRATEGY_COLORS (routing-strategy hues), LAYER_COLORS (compression layer pills), the provider brand color in ProviderTopology, and IoNode's indigo/green input-output identity pair. The new test asserts both halves — what became a token AND what stays hex — so a later sweep cannot silently swallow a categorical palette. Refs diegosouzapw#12378
…egosouzapw#12876) Validado numa worktree combinada com a onda de dashboard/monitoring desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 130/131 nos testes focados — a falha restante é asserção de tempo de parede sob carga, verde 6/6 isolada. Um health que diz "falhou" sem dizer **qual** conexão obriga o operador a cruzar logs para achar o óbvio. Expor os ids das que falharam é o que transforma o endpoint em ferramenta de diagnóstico.
…iegosouzapw#12882) Validado numa worktree combinada com a onda de dashboard/monitoring desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 130/131 nos testes focados — a falha restante é asserção de tempo de parede sob carga, verde 6/6 isolada. Paginar e completar o stream em `GET /v1/models` é a correção certa para catálogo grande: um payload único que cresce com o número de providers vira timeout silencioso no cliente, não erro.
…n on one model 402 (diegosouzapw#12875) Validado numa worktree combinada com a onda de dashboard/monitoring desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 130/131 nos testes focados — a falha restante é asserção de tempo de parede sob carga, verde 6/6 isolada. Envenenar a conexão inteira por um 402 de **um** modelo é o erro clássico de granularidade em provider openai-compatible com múltiplos upstreams — derruba modelos que estavam saudáveis. Restringir ao modelo afetado é o comportamento correto, e é a mesma distinção que o guia de resiliência faz entre cooldown de conexão e lockout de modelo.
…instead of "(empty)" (diegosouzapw#12727) Validado numa worktree combinada com a onda de dashboard/monitoring desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 130/131 nos testes focados — a falha restante é asserção de tempo de parede sob carga, verde 6/6 isolada. "(empty)" para um nó de ferramenta ainda não resolvido é informação errada, não ausência de informação — o usuário lê como "não retornou nada". Spinner de pendente diz a verdade.
…egosouzapw#12937) Validado numa worktree combinada com a onda de dashboard/monitoring desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 130/131 nos testes focados — a falha restante é asserção de tempo de parede sob carga, verde 6/6 isolada. Sentinela que não se explica (`—`, `?`) faz o leitor inventar a razão. Explicar no hover é metade; o `check:radar-sentinels` é a outra — sem o gate, a explicação apodrece na primeira coluna nova.
…iegosouzapw#12884) Um `ALL_TARGETS_SKIPPED` 503 que não diz qual janela esgotou é opaco justamente no momento em que o operador mais precisa saber. Alinhar os rótulos de janela AUTH com os da API de uso fecha a outra metade: dois nomes para a mesma coisa fazem o dashboard e o erro parecerem discordar. Revalidei sobre o tip: **6/6**, typecheck:core limpo, check-file-size OK. **Dois consertos meus na sua branch.** 1. `typecheck:core` falhava com `TS2345` em `comboAttemptLoop.ts` (linhas 130 e 416): o `QuotaSkipTarget` declarava `connectionId?: string`, mas o `ResolvedComboTarget` carrega `string | null` para alvo não-pinado. Alarguei para `string | null` no tipo de diagnóstico em vez de estreitar o call site — o módulo só **lê** o campo e a linha 29 já narrowa com `typeof === "string"`, então null não custa nada ali. Isso apareceu porque o `comboAttemptLoop` mudou de forma no diegosouzapw#12746/diegosouzapw#12811, mergeados nesta mesma campanha depois que você cortou a branch. 2. O `roundRobinCombo.ts` foi de 1198 para 1205 e cruzou o teto de 1200 para arquivo novo. Congelei com justificativa: o arquivo já nasceu em 1198 quando o diegosouzapw#12811 o levantou de dentro do `combo.ts`, e os diagnósticos em si vivem no `quotaSkipDiagnostics.ts`, sob o cap. Registrei que a próxima extração natural é o corpo do attempt loop, mas que ele acabou de ser movido e deve assentar antes de ser cortado de novo.
diegosouzapw
merged commit Sep 10, 2026
a152eb9
into
diegosouzapw:release/v3.8.51
4 of 7 checks passed
thinh0704hcm
added a commit
to thinh0704hcm/OmniRoute
that referenced
this pull request
Sep 12, 2026
Brings canonical up to upstream release/v3.8.51 (51 commits), including the queue-bound resilience fix a152eb9 (diegosouzapw#12715). Conflicts resolved: - open-sse/services/rateLimitManager.ts: took upstream's per-connection queue budget (maxWaitMs shared across the provider gate, default slot and the Bottleneck queue) and re-applied the fork overlay's linked abort signal on top, so an expired execution backstop also aborts the in-flight upstream request instead of leaking it and holding a slot queued work must wait on. - open-sse/utils/stream.ts: took upstream's buildUsageOnlyChunk(id, model, usage) call shape; the fork's stale object form set `choices: []` as the usage payload.
5 tasks
diegosouzapw
added a commit
to initguru/OmniRoute
that referenced
this pull request
Sep 15, 2026
…t gate behavior Answers the open technical question from diegosouzapw#12902's review: does a GLOBAL maxWaitMs=0 reintroduce the unbounded-queue regression diegosouzapw#12715 fixed (a request hanging ~6min until the client aborts)? Evidence, exercising the real gate chatCore.ts actually calls (accountSemaphore.acquireMany({ timeoutMs: requestQueue.maxWaitMs }), not the Bottleneck reservoir the PR's own tests cover) under real contention (maxConcurrency=1, two concurrent acquires): - No: it does not hang. setTimeout(reject, 0) fires on the next tick, so a second contending request is rejected with SEMAPHORE_TIMEOUT in low milliseconds, never minutes. - But it is also not a genuine 'no cap' — an operator setting 0 expecting 'wait as long as it takes' instead gets near-zero tolerance for even momentary contention on any configured concurrency gate (global/provider/account). This is a real asymmetry vs. the Bottleneck reservoir path (where 0 truly means unbounded) left for the maintainer to decide how to resolve — not something this pass can decide unilaterally. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
diegosouzapw
added a commit
that referenced
this pull request
Sep 17, 2026
…expiration (#12902) * fix(resilience): allow maxWaitMs=0 as disable sentinel for execution expiration maxWaitMs normalization clamped the value to min:1, silently rewriting an operator's 0 ("disable the limiter-managed execution deadline") into 1 — a 1ms expiration that killed every long-running job instantly. This broke long-running reasoning models (GLM-5.2 with reasoning.effort=max spends minutes before the first token, exceeding any practical maxWaitMs; the TTB safety net is FETCH_TIMEOUT_MS, default 600s). Fix: lower the floor to min:0 so 0 is preserved as the disable sentinel. Issue #4165 follow-up. Tests: 7/7 (resilience-normalize-maxwaitms-disable 5 + rate-limit- maxwaitms-disable-execution 2). typecheck:core clean. * fix(resilience): relax requestQueueSettingsSchema.maxWaitMs to allow 0 normalizeRequestQueueSettings already treats maxWaitMs=0 as an explicit disable sentinel (queue-wait budget off), but the settings API schema still rejected 0 with min(1), so an operator could never actually reach the fix through PATCH /api/resilience. executionMaxWaitMs is untouched (stays min(1) — separate field, separate decision, see #12902 item 4). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * test(resilience): prove maxWaitMs=0 vs #12715's queue-wait gate behavior Answers the open technical question from #12902's review: does a GLOBAL maxWaitMs=0 reintroduce the unbounded-queue regression #12715 fixed (a request hanging ~6min until the client aborts)? Evidence, exercising the real gate chatCore.ts actually calls (accountSemaphore.acquireMany({ timeoutMs: requestQueue.maxWaitMs }), not the Bottleneck reservoir the PR's own tests cover) under real contention (maxConcurrency=1, two concurrent acquires): - No: it does not hang. setTimeout(reject, 0) fires on the next tick, so a second contending request is rejected with SEMAPHORE_TIMEOUT in low milliseconds, never minutes. - But it is also not a genuine 'no cap' — an operator setting 0 expecting 'wait as long as it takes' instead gets near-zero tolerance for even momentary contention on any configured concurrency gate (global/provider/account). This is a real asymmetry vs. the Bottleneck reservoir path (where 0 truly means unbounded) left for the maintainer to decide how to resolve — not something this pass can decide unilaterally. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Jihyun Son <jihyun.son@sk.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
diegosouzapw
added a commit
that referenced
this pull request
Sep 19, 2026
…ests + pack-policy + dashboard-typecheck) Three production defects the tests caught: - rateLimitManager: maxWaitMs=0 (the #12902 disable sentinel) hit #12715's queue-budget gate as "0 ms left" and 503'd every protected request. - emergencyFallback: #14006 silently switched the budget-exhaustion target provider nvidia -> groq against ENVIRONMENT.md and the NIM snapshot; restored. - claudeConnectionFields.ts vs ClaudeConnectionFields.tsx (#13074) differed only by casing; helpers renamed to claudeConnectionFieldValues.ts. Guards realigned to legitimate changes: #13874 rotation map (distinct token in the error test), #13350 origin-IP denylist, #13318 shared-catalog growth (counts by invariant), comboTargetKeyPolicy import in the telegram stub, the 22 README mirrors that #13940/#14106 stamped with the retired openference.svg (translated Cerebras cells recovered from history, hashes re-stamped), bin/antigravity-bridge.mjs allowed in the pack policy, and the two dashboard typecheck regressions (typed pinned section, ComponentProps cast). Refs #13866.
diegosouzapw
added a commit
that referenced
this pull request
Sep 21, 2026
#12902 released requestQueue.maxWaitMs=0 as the sentinel that disables the queue-wait deadline, but the #12715 queue-budget gate in withRateLimit() (`if (queueRemainingMs <= 0) throw`) read 0 as "budget spent" and rejected every request on a protected connection with an immediate 503 queue-budget error — the exact opposite of what the setting promises. rate-limit-maxwaitms-disable-execution ("400ms job completes without 504") was red on the tip. When no caller budget is passed and the configured queue budget is 0, skip the gate, never arm the queue-wait timer and hand awaitProviderDefaultSlot no budget (it falls back to the window). Execution stays bounded by executionMaxWaitMs and the upstream fetch-start timeout, as before. Refs #13866
diegosouzapw
added a commit
that referenced
this pull request
Sep 21, 2026
Drains base-red waves 5–8 of release/v3.8.51: 35+ tests and the pack-policy, api-typecheck, dashboard-typecheck, docs-all, agent-skills-sync, ESLint and mutation-coverage gates (#13866). Production defects the tests caught: - Caveman: #12825's file-pack prefilter tested anchored rules against the original text, so leader_phrases never ran. - pack-artifact: httpClientAbortGuard.mjs was missing from the staging allowlist and the required set — every published boot died with ERR_MODULE_NOT_FOUND (#14191). - rateLimitManager: maxWaitMs=0 (the #12902 disable sentinel) hit #12715's queue-budget gate as "0 ms left" and 503'd every protected request. - emergencyFallback: #14006 silently switched nvidia -> groq; restored per ENVIRONMENT.md and the NIM snapshot. - ClaudeConnectionFields.tsx vs claudeConnectionFields.ts (#13074) differed only by casing; helpers renamed to claudeConnectionFieldValues.ts. - /v1/responses/input_tokens body validated with Zod (HR#7). Guards realigned to legitimate changes (#12663, #13863, #12565, #13990, #13874, #13350, #13318, #13848 typed and split out of a size-capped file — #14254), i18n catalogs for the 7 keys of #13074/#7f1b4a5e in 65 locales, 22 README mirrors restored to the Cerebras cell, env docs for 5 vars, regenerated omni-version-manager skill. Refs #13866. Closes #14254. Refs #14191. Co-authored-by: Prabhjot Singh <jotgill1522@gmail.com> Co-authored-by: Xmon Dai <xiechimon@qq.com>
diegosouzapw
added a commit
that referenced
this pull request
Sep 22, 2026
…ll, TS2677, ESLint (Refs #13866) (#14331) * fix(build): ship httpClientAbortGuard.mjs in the pack artifact; validate input_tokens with Zod Wave five of the release/v3.8.51 base-reds, part 1 — the two that matter. #14064 restored server-ws.mjs's import of ./httpClientAbortGuard.mjs and the assembleStandalone copy, but not the two pack-artifact policy entries that were lost with it. Without APP_STAGING_ALLOWED_EXACT_PATHS the prepublish prune deletes the file; without PACK_ARTIFACT_REQUIRED_PATHS nothing notices. Every boot of the published package would die with ERR_MODULE_NOT_FOUND — the 3.8.47 head-response-guard class. Both closure suites (9/9) now enforce it. #13910's /v1/responses/input_tokens read request.json() behind a hand-rolled typeof check. Hard Rule #7 wants the boundary on Zod; the t06 guard caught it. Same passthrough envelope the catch-all Responses route uses, since the counters below already walk the fields defensively. 9/9 on the route's suite. Five no-unused-vars left behind by the wave (cliRuntime execFileSync, arena test symbols and a type, compression rmSync, waitForServer req) are removed. The 'openwa routes removed without deprecation' entry from the #14101 run was an artifact of that PR trailing its base — the gate is clean on the tip. Refs #13866 * fix(compression): let anchored file-pack rules see the transformed text; align wave-5 guards Wave five of the release/v3.8.51 base-reds, part 2. One production defect. #12825 (Hungarian Caveman pack) stopped gating file-pack rules with the English keyword list and tested the rule's own regex instead — against `lowerResult`, a lower-cased copy of the ORIGINAL text that the loop never refreshed. An anchored pattern like leader_phrases' `^(?:i will|…)` therefore ran its prefilter on "sure, i will…", failed the anchor, and was skipped; the rule that strips "I will " from every English response was dead since the merge. The prefilter now sees the text as the rules so far have left it. New test fails on the tip and passes here; all Caveman suites, Hungarian included, are 95/95. A frozen no-unused-vars suppression on caveman.ts no longer had a target and is pruned. Two more TS2677 predicates of the kind #14101 fixed: #13910 (rerankProviderNodes.ts, `n is RerankProviderNodeRow` on a Record row) and #13957's mitm catalog (antigravity.ts, `c is DynamicCatalogModel` on a literal-or-null). Both narrow by NonNullable of the element's own type; the api-route typecheck was 285 against a baseline of 283 on the pristine tip. The rest are guards trailing legitimate changes: - #12663 made gemini-3.8-flash the catalog head; T28 pinned 3.7. - #13863 put mimo-v2.5 into the shared vision heuristic on purpose (the base model is multimodal, only the Pro variants are text-only). The safety test now asserts the real invariant: base and :free aliases yes, -pro no. - #12565 moved npm-prefix detection into cliRuntimeNpmPrefix.ts with a process-lifetime cache that importFresh() does not reset; the case resets it. #12565 also builds Windows candidates with path.win32 on purpose; the qodercli test compared against POSIX path.join. - #13990 (the 2 GB Docker image) copies better-sqlite3 with --chown; the guard matched the flag order literally. Now flag-order tolerant, still fails when --from=builder is removed. - #13378 reintroduced public/openference.svg under a name #11750 retired for missing provenance and swapped the Cerebras showcase cell for it. The cell is back and the asset is gone; whether the new drawing counts as provenance is the owner's call. Refs #13866 * fix(i18n): translate the 7 sidebar-pin and Claude low-priority keys into all 65 locales #7f1b4a5e (sidebar pinned items) and #1b2349de (Claude OAuth lower-priority / auto-reset) landed with their 7 new keys in en.json only, which the vi and pt-BR parity suites flag. Translated with the repo's own sync-ui-keys --translate-markers against the .113 i18n instance (codex/gpt-5.6-sol-low): +446 lines across 65 catalogs, zero __MISSING__ markers, placeholders intact. vi.json also has two keys reordered to mirror en.json; values unchanged. Refs #13866 * chore(quality): list the 8 covering tests the sixth wave added in stryker tap.testFiles 30 commits landed on release/v3.8.51 while wave five drained; eight new unit tests cover mutated modules and were not in tap.testFiles, so their mutant kills did not count and check:mutation-test-coverage --strict failed on the merged tree. Appended at the end of the list, nothing reordered. Refs #13866 * docs: document the five env vars of the 09-18 wave; regenerate the version-manager skill for the open-wa routes check:docs-all: BRIDGE_PORT, ROUTER_URL and CERT_DIR (bin/antigravity-bridge.mjs, #c74cea3d), OPENWA_SERVICE_PORT (src/lib/services/bootstrap.ts, #1e8c913c) and NEXT_PUBLIC_PORT (src/shared/hooks/useDisplayBaseUrl.ts, #d715190b) were read in code but absent from .env.example and docs/reference/ENVIRONMENT.md. Added next to their neighbours, with the defaults the code actually uses (open-wa is 8323, not the 201xx range the other services sit in). check:agent-skills-sync: the open-wa feature added eight /api/services/openwa/* routes to docs/openapi.yaml without regenerating skills/omni-version-manager/ SKILL.md. Regenerated with the repo generator; the diff is exactly those eight route sections. Refs #13866 * test: register the crash guard in the pack snapshot; inventory #13874's refresh-lane row read pack-artifact-policy pins the list of root runtime files check:pack-artifact must find in the tarball; dist/httpClientAbortGuard.mjs joined PACK_ARTIFACT_REQUIRED_PATHS in this PR and the snapshot follows. #13874 re-reads the connection row inside the Claude refresh lane so a queued health check does not POST a refresh token a Layer 2 refresh already rotated — a state read, inventoried like the family-cooldown lookup (tokenHealthCheck.ts 2 -> 3). Refs #13866 * chore(quality): list native-codex-auto-resume test in stryker tap.testFiles (#13180 landed without it) * fix(release): drain the seventh base-red wave of release/v3.8.51 (9 tests + pack-policy + dashboard-typecheck) Three production defects the tests caught: - rateLimitManager: maxWaitMs=0 (the #12902 disable sentinel) hit #12715's queue-budget gate as "0 ms left" and 503'd every protected request. - emergencyFallback: #14006 silently switched the budget-exhaustion target provider nvidia -> groq against ENVIRONMENT.md and the NIM snapshot; restored. - claudeConnectionFields.ts vs ClaudeConnectionFields.tsx (#13074) differed only by casing; helpers renamed to claudeConnectionFieldValues.ts. Guards realigned to legitimate changes: #13874 rotation map (distinct token in the error test), #13350 origin-IP denylist, #13318 shared-catalog growth (counts by invariant), comboTargetKeyPolicy import in the telegram stub, the 22 README mirrors that #13940/#14106 stamped with the retired openference.svg (translated Cerebras cells recovered from history, hashes re-stamped), bin/antigravity-bridge.mjs allowed in the pack policy, and the two dashboard typecheck regressions (typed pinned section, ComponentProps cast). Refs #13866. * test: type the #13848 Gemini pairing tests (no-explicit-any) and inventory the semantic-cache embedding picker's connection read Both arrived with the tip merge: #13848 added 13 explicit any casts to translator-openai-to-gemini.test.ts (no-explicit-any is an error under tests/), and 7a92129's embeddingOptions.ts reads provider connections once without a hard-session-lease inventory entry. Stale suppression count pruned for the test file only. Refs #13866. * test: split the #13848 turn-pairing cases out of translator-openai-to-gemini.test.ts The file sits exactly at its frozen size cap; typing the pairing tests (no-explicit-any) pushed it 14 lines over. The two cases are a coherent regression suite of their own, so they move to translator-openai-to-gemini-turn-pairing-13848.test.ts (registered in stryker tap.testFiles) instead of widening the baseline. * docs(env): document BRIDGE_PORT, ROUTER_URL, CERT_DIR, OPENWA_SERVICE_PORT and NEXT_PUBLIC_PORT (Refs #13866) check:env-doc-sync has been red on the release tip since these five vars reached code without their .env.example / ENVIRONMENT.md entries: bin/antigravity-bridge.mjs (BRIDGE_PORT, ROUTER_URL, CERT_DIR — #14006), src/lib/services/bootstrap.ts + api/services/openwa/_lib.ts (OPENWA_SERVICE_PORT) and src/shared/hooks/useDisplayBaseUrl.ts (NEXT_PUBLIC_PORT — #13533). Defaults and source files copied from the reads themselves. * chore(skills): regenerate omni-version-manager for the open-wa service routes (Refs #13866) check:agent-skills-sync (Merge integrity job) has been red on the tip since the open-wa embedded-service routes reached docs/openapi.yaml without the generated SKILL.md being refreshed. Output of scripts/skills/generate-agent-skills.mjs --apply, no hand edits: the eight /api/services/openwa/* operations. * fix(types): make the two TS2677 type predicates sound (Refs #13866) check:api-typecheck has been red on the tip with two "type predicate's type must be assignable to its parameter's type" errors: - src/app/api/v1/_shared/rerankProviderNodes.ts (#13733): the read cache hands back `Record<string, unknown> | null`, and an interface whose members are all optional is not assignable to an index-signature type. Narrow to the non-null record and assert the row shape afterwards. - src/mitm/handlers/antigravity.ts (#14006): the map callback returned `{ displayName: string }` while DynamicCatalogModel declares it optional, so the predicate could not be proven. Type the callback's return explicitly and filter on `!== null`. No runtime change; rerank-remote-provider-nodes / rerank-local-node-shapes / mitm-handler-antigravity stay green. * fix(lint): clear the 92 ESLint errors the lint gate reports on the tip (Refs #13866) - tests/unit/translator-openai-to-gemini.test.ts: #13848 / #13318 added 13 `any` casts/params on top of the 74 frozen for the file, so ESLint reported all 87. Typed them (GeminiRequestWithContents / GeminiToolPart, and the existing GeminiRequestWithConfig) and pruned the file's suppression to the new count of 71 — nothing else in eslint-suppressions.json changes. - no-unused-vars: execFileSync import (src/shared/services/cliRuntime.ts, #12565), getArenaEloSyncStatus + makeLeaderboardMap + ArenaLeaderboardMap (tests/unit/arena-elo-sync-redesign.test.ts, #13446), rmSync (compressionAnalyticsWriterFlatRate.test.ts, #13446), `req` → `_req` (waitForServer-slow-first-response.test.mjs). translator-openai-to-gemini 48/48; arena-elo-sync-redesign, compressionAnalyticsWriterFlatRate, waitForServer-slow-first-response green. * fix(compression): stop skipping anchored Caveman rules that only match after earlier rules #12825 (Hungarian pack) replaced the English keyword prefilter with a `rule.pattern.test(lowerText)` pre-check for every file-based rule, including the default `en` pack. `lowerText` is the ORIGINAL message, so anchored rules such as `leader_phrases` (`^i will …`) — which only match after `pleasantries` strips "Sure, " — were dropped before they could run. `caveman-v379` caught the regression ("I will ensure …" survived at full intensity). Tag file-based rules with their pack language in ruleLoader and let the keyword prefilter apply to `en`/built-in rules only; non-English packs (which reuse English rule names) simply run their localized regex, which is what the pre-test cost anyway. Drops the now-unused CAVEMAN_RULES import and prunes the already-stale `caveman.ts` no-unused-vars suppression (0 violations on the tip) that blocked the pre-commit hook for any change to this file. Refs #13866 * test(models): align catalog and vision-heuristic guards with the tip's intended contracts Three base-reds where the production change was deliberate and the pinned guard was simply not bumped by the PR that changed the contract: - agy-antigravity-shared-catalog-12724: #13318 added the three Gemini 3.8 Flash tiers (high/medium/low, no "-tiered" endpoint for 3.8) to the shared Antigravity/AGY base, 10 -> 13. Pin the new size in one constant and make the buildSurfaceCatalog delta assertions relative to it. - t28-model-catalog-updates: #12663 (issue #12638) registered gemini-3.8-flash at the head of the AI Studio fallback catalog as the current Flash default; assert 3.8 first and keep 3.7 present. - command-code-mimo-v2-5-safety: #13863 (issue #13847) added an explicit "mimo-v2.5" fragment to the shared vision heuristic so provider-qualified and `-free` aliases keep their vision flag. The guard's real concern (the "mimo-vl" fragment must not cover "mimo-v2.5") is asserted on the fragment itself; the bare id is now vision by heuristic on purpose, and the Pro text-only sibling stays excluded. Refs #13866 * test(cli): follow the #12565 cliRuntime module split in the npm-prefix and qodercli guards #12565 (issue #12563) moved the npm global-prefix cache out of cliRuntime.ts into cliRuntimeNpmPrefix.ts and built the Windows known-bin candidates with `path.win32` (cliRuntimeWindowsNode.ts) so they stay Windows-shaped when `process.platform` is mocked on a POSIX runner. Two pre-existing guards depended on the old layout: - cli-runtime-extended "resolves known binaries from npm global prefix": importFresh() only re-evaluates cliRuntime.ts; the prefix cache now lives in a module that stays shared across cases, so a real `npm config get prefix` from an earlier case was cached and the mocked execFileSync never ran. Reset the cache with the helper #12565 exported for exactly this in afterEach. - qodercli-windows-resolve-6263: compare against `path.win32.join` — identical to `path.join` on a real Windows host, which is the behaviour under test. Production behaviour is unchanged on both platforms. Refs #13866 * test(auto-update): write the source-mode log inside the test's own temp dir The launchAutoUpdate case pointed AUTO_UPDATE_LOG_PATH at a fixed, world-shared `/tmp/auto-update-source.log`. On the .113 runner the suite executes both as `root` and as `runner` (uid 1001): the file survives owned by whoever ran first (`-rw-r--r-- root root`), and the next `openSync(logPath, "a")` fails with EACCES for the other user. Reproduced locally by making the shared file read-only; production code is untouched (autoUpdate.ts last changed in #9354). Use a per-test mkdtemp path for the source-mode log and clean the whole temp root in the existing finally block. Refs #13866 * fix(dashboard): rename claudeConnectionFields.ts so it no longer case-collides with ClaudeConnectionFields.tsx #13074 added two modules to the provider-detail modals directory whose names differ only by casing: `ClaudeConnectionFields.tsx` (the component) and `claudeConnectionFields.ts` (the value/patch helpers). On a case-insensitive filesystem the pair breaks the webpack build (#6584 guard), and esbuild's resolver already picks the `.tsx` for the extension-less `./claudeConnectionFields` specifier, so the provider-detail client entry failed to bundle ("No matching export ... for import claudeConnectionFieldPatch"). Rename the helper module to `claudeConnectionFieldValues.ts` (the same naming the sibling `quotaScrapingFieldValues.ts` uses) and point the only importer, EditConnectionModal.tsx, at the new name. Greens tests/unit/case-collision-6584.test.ts and tests/unit/media-page-client-browser-bundle.test.ts. Refs #13866 * fix(build): allowlist dist/httpClientAbortGuard.mjs so the published tarball keeps the server-ws crash guard #14064 (re-land of #13636) made scripts/dev/standalone-server-ws.mjs import ./httpClientAbortGuard.mjs and taught assembleStandalone to copy the shared implementation next to dist/server-ws.mjs — but never registered the file in scripts/build/pack-artifact-policy.ts. The prepublish prune deletes anything outside APP_STAGING_ALLOWED_EXACT_PATHS, and check:pack-artifact only fails on PACK_ARTIFACT_REQUIRED_PATHS entries, so the next `omniroute` tarball would boot straight into ERR_MODULE_NOT_FOUND (the #7065 / tls-options class the closure tests exist to catch). Add the bare and dist/ entries to both lists and extend the required-paths snapshot in tests/unit/pack-artifact-policy.test.ts. Greens tests/unit/pack-artifact-entrypoint-closures.test.ts and tests/unit/pack-artifact-server-ws-closure.test.ts. Refs #13866 * test(docker): accept --chown=node:node on the better-sqlite3 runner COPY #14010 deliberately changed the runner-stage COPYs to `COPY --chown=node:node --from=builder ...` (ownership at copy time instead of a second ~2 GB `chown -R` overlay layer). The Dockerfile contract test still matched the old `COPY --from=builder /app/node_modules/better-sqlite3` prefix and went red on the tip even though the native-addon guard it protects is intact. Tolerate the optional --chown flag; every other assertion (node-gyp rebuild, both `test -f .../better_sqlite3.node` checks) is unchanged. Refs #13866 * fix(api): validate /v1/responses/input_tokens bodies with Zod (t06) #13167 added the local Responses token-count route with hand-rolled `typeof` checks on `request.json()`. Hard Rule #7 and the t06 gate (scripts/check/check-route-validation.mjs, mirrored by tests/unit/route-body-validation-t06.test.ts) require every route that reads request.json() to go through validateBody()/safeParse(), so the tip was red. Add `v1ResponsesInputTokensSchema` (pins the wire types the counter reads — model/instructions strings, input string-or-array, tools array — and lets unknown keys through since they are counted, never forwarded) and run the body through validateBody(); a type mismatch is now a 400 naming the field instead of a silently ignored key. Regression test added to tests/unit/responses-input-tokens-local-route.test.ts. Refs #13866 * fix(docs): drop the retired openference.svg asset reintroduced by #13378 `openference.svg` is one of the 78 provider assets retired for missing provenance (tests/unit/provider-assets-generic-fallback.test.mjs freezes that list and forbids any tracked surface from referencing a retired name). #13378 added a new hand-drawn `public/openference.svg` outside the manifest-audited public/providers/ tree and pointed the README free-tier table (plus the 22 i18n mirrors that carry the row) at it, which put the retired name back on a tracked surface and left an unaudited asset in the package. Use the generic fallback icon (`public/providers/cli-generic.svg`) the other provenance-less providers already use, delete the unaudited file, and adopt the mechanical README edit into .i18n-state.json (`i18n:run -- --adopt --files=README.md`, no API calls) so the i18n drift gate does not flag README.md as source-changed. Refs #13866 * test(lease): classify the two connection-query sites added by #14159 and #13874 The hard-lease bypass inventory froze every getProviderConnections / getProviderConnectionById site with a class; two landed on the tip without a golden update: - src/app/api/settings/cache-config/embeddingOptions.ts (#14159, re-land of #12630): read-only listing that feeds the semantic-cache embedding dropdown, same shape as the qdrant embedding-models route — class C. - src/lib/tokenHealthCheck.ts 2 -> 3 (#13874): re-reads the row by id after an unrecoverable refresh error to detect credentials rotated by a concurrent Layer 2 refresh before deactivating — a state read, not dispatch; stays C. Refs #13866 * fix(sse): restore nvidia as the emergency budget-fallback provider #14006 (Antigravity MITM catalog injection) flipped EMERGENCY_FALLBACK_CONFIG.provider from "nvidia" to "groq" in one line, without touching ENVIRONMENT.md, .env.example, the chat.ts comment or the NVIDIA hosted-model snapshot, all of which still promise nvidia/openai/gpt-oss-120b. Operators without a Groq connection got the original 402 back instead of the free reroute, and chat-route-coverage ("uses the emergency fallback model on budget exhaustion" / "returns the primary budget error when emergency fallback also fails") went red on the tip. Put the documented default back; the #14006 bridge tests exercise bin/antigravity-bridge.mjs and do not read this config. Refs #13866 * fix(resilience): keep maxWaitMs=0 a "no queue deadline" sentinel #12902 released requestQueue.maxWaitMs=0 as the sentinel that disables the queue-wait deadline, but the #12715 queue-budget gate in withRateLimit() (`if (queueRemainingMs <= 0) throw`) read 0 as "budget spent" and rejected every request on a protected connection with an immediate 503 queue-budget error — the exact opposite of what the setting promises. rate-limit-maxwaitms-disable-execution ("400ms job completes without 504") was red on the tip. When no caller budget is passed and the configured queue budget is 0, skip the gate, never arm the queue-wait timer and hand awaitProviderDefaultSlot no budget (it falls back to the window). Execution stays bounded by executionMaxWaitMs and the upstream fetch-start timeout, as before. Refs #13866 * test: align three fixtures with the #13874, #13861 and #13350 contracts Three base-reds that are deliberate contract changes, not defects: - executor-default-base "refreshCredentials swallows refresh errors": #13874 records rotations on the Layer 2 (no connectionId) refresh path too, so the "refresh-me" token the previous case already rotated was served from the rotation map without the network POST the test wanted to fail. Use a token nobody rotated. - telegram-keycache-bounded-13165: #13861 made comboTargetKeyPolicy import isModelBlockedByPatterns from db/apiKeys; the loader-stubbed module lacked it and the suite died at module load. Export an honest "not blocked" stub (the test has no blocked models). - upstream-headers-proxy-auth "ordinary headers are still allowed": #13350 forbids the whole origin-IP forwarding set upstream (covered by upstream-headers-sanitize). Swap x-forwarded-for for x-request-id. Refs #13866
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…iegosouzapw#12715) Fila sem teto que segura a request seis minutos até o cliente abortar é pior que 503 imediato: consome slot, mascara a saturação e ainda entrega erro no fim. Um orçamento `maxWaitMs` por conexão compartilhado entre gate, slot padrão do provider e fila do Bottleneck é a forma certa — o teto tem que ser um só, senão cada camada espera o seu. O `max(perConn, upstream)` no `executionMaxWaitMs` é o detalhe que evita a correção matar request em voo, que seria trocar um defeito por outro. Registro a atribuição: você manteve o diegosouzapw#12635 aberto para o @Tushar49 e creditou a percepção dele (providers lentos precisam de 2min→10min por conexão) enquanto adiciona o encanamento que faltava. É o jeito certo de construir sobre PR de outra pessoa sem tomar o crédito. Sobre o `npm run lint` desmarcado com a nota do eslint quebrado no ambiente: deixar em branco e explicar vale mais que marcar sem ter rodado. Rodei aqui: limpo. Revalidei sobre o tip: **13/13**, typecheck:core limpo, check-file-size OK. O `file-size-baseline.json` conflitou com os rebaselines desta campanha — resolvido aditivamente, JSON revalidado com `json.load`.
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…expiration (diegosouzapw#12902) * fix(resilience): allow maxWaitMs=0 as disable sentinel for execution expiration maxWaitMs normalization clamped the value to min:1, silently rewriting an operator's 0 ("disable the limiter-managed execution deadline") into 1 — a 1ms expiration that killed every long-running job instantly. This broke long-running reasoning models (GLM-5.2 with reasoning.effort=max spends minutes before the first token, exceeding any practical maxWaitMs; the TTB safety net is FETCH_TIMEOUT_MS, default 600s). Fix: lower the floor to min:0 so 0 is preserved as the disable sentinel. Issue diegosouzapw#4165 follow-up. Tests: 7/7 (resilience-normalize-maxwaitms-disable 5 + rate-limit- maxwaitms-disable-execution 2). typecheck:core clean. * fix(resilience): relax requestQueueSettingsSchema.maxWaitMs to allow 0 normalizeRequestQueueSettings already treats maxWaitMs=0 as an explicit disable sentinel (queue-wait budget off), but the settings API schema still rejected 0 with min(1), so an operator could never actually reach the fix through PATCH /api/resilience. executionMaxWaitMs is untouched (stays min(1) — separate field, separate decision, see diegosouzapw#12902 item 4). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * test(resilience): prove maxWaitMs=0 vs diegosouzapw#12715's queue-wait gate behavior Answers the open technical question from diegosouzapw#12902's review: does a GLOBAL maxWaitMs=0 reintroduce the unbounded-queue regression diegosouzapw#12715 fixed (a request hanging ~6min until the client aborts)? Evidence, exercising the real gate chatCore.ts actually calls (accountSemaphore.acquireMany({ timeoutMs: requestQueue.maxWaitMs }), not the Bottleneck reservoir the PR's own tests cover) under real contention (maxConcurrency=1, two concurrent acquires): - No: it does not hang. setTimeout(reject, 0) fires on the next tick, so a second contending request is rejected with SEMAPHORE_TIMEOUT in low milliseconds, never minutes. - But it is also not a genuine 'no cap' — an operator setting 0 expecting 'wait as long as it takes' instead gets near-zero tolerance for even momentary contention on any configured concurrency gate (global/provider/account). This is a real asymmetry vs. the Bottleneck reservoir path (where 0 truly means unbounded) left for the maintainer to decide how to resolve — not something this pass can decide unilaterally. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Jihyun Son <jihyun.son@sk.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
fix(resilience): stop unbounded queue that hangs 6min until Aborted— one per-connectionmaxWaitMsbudget shared across gate + provider-default slot + Bottleneck queue (fail-closed 503);executionMaxWaitMsper-connection withmax(perConn, upstream)so it never kills a mid-flight request. Builds on the need surfaced in feat(rate-limit): raise per-connection maxWaitMs cap from 2min to 10m… #12635 — keeps @Tushar49's insight (slow providers need 2min→10min per-connection headroom) and adds the missing plumbing (#12635blocked byh "audio-transcriptions"typo; a ceiling bump alone does not bound the wait). feat(rate-limit): raise per-connection maxWaitMs cap from 2min to 10m… #12635 stays open for its author (with thanks to @Tushar49).Related Issues
Validation
npm run lint— targetedeslinton the touched files is clean; the full run is red on the base (🔴 Release branch not green: release/v3.8.51 #12732), so the box stays unchecked.Focused suites, all green on this branch after the change:
rate-limit-remaining-budget(5/5),rate-limit-manager-queue-bound(4/4),rate-limit-execution-per-conn(4/4), plus therate-limit-execution-timeout-message-4165+provider-rate-limit-overrides-schemacompanions (25/25). No VPS gate applies — it's local queue/budget plumbing exercised by unit tests, no upstream round-trip involved.Tests Added Or Updated
tests/unit/rate-limit-remaining-budget.test.ts— shared budget gate+slot+queue (remaining propagates, 503 on exhaust).tests/unit/rate-limit-manager-queue-bound.test.ts— bounded Bottleneck queue with anti-orphan guardremaining<=0not queued, abort vs timeout distinct).tests/unit/rate-limit-execution-per-conn.test.ts— per-connectionexecutionMaxWaitMsand upstream clamp.Coverage Notes
Touches
open-sse/handlers/chatCore.ts,open-sse/handlers/chatCore/queueBudget.ts(new),config/quality/file-size-baseline.json,open-sse/services/rateLimitManager.ts,open-sse/services/providerDefaultRateLimit.ts,src/lib/resilience/settings/types.ts,src/shared/validation/schemas/provider.ts.Covered by the 3 new suites plus
rate-limit-execution-timeout-message-4165.test.tsandprovider-rate-limit-overrides-schema.test.ts. No file loses coverage.Reviewer Notes
503on budget exhaust) vs queue depth (429onSEMAPHORE_QUEUE_FULL;504on execution expiration) intentionally distinct — only bare time-budget timeouts are shaped to legacy 503.executionMaxWaitMsismax(perConn, upstream)viaopts:{executor,providerSpecificData}passed fromhandleChatCoretowithRateLimit.chatCore.tsis a frozen file already at its ceiling. The admission error shaping lives inopen-sse/handlers/chatCore/queueBudget.ts; the remaining call-site wiring (+15 lines) is rebaselined inconfig/quality/file-size-baseline.jsonwith a note, like the existing_rebaseline_*entries. The gate-ordering contractchatcore-hierarchical-admissionpasses its first case; its second case fails the same way on the base tip.