feat(dashboard): badge a conversation that never reached a clean stop - #12717
Merged
diegosouzapw merged 79 commits intoSep 10, 2026
Merged
diegosouzapw merged 79 commits into
diegosouzapw merged 79 commits into
Conversation
Live incident: a reasoning-heavy stream blew past createStructuredSSECollector's own event-count cap mid-stream (3013 SSE events, 1487 dropped). The stored clientResponse never reached "completed" (_truncated: true, summary.status stuck at "in_progress", output: []) -- the client never got a valid reply, and nothing on /dashboard/conversations showed this had happened. The conversation just sat there looking like any other row. A bare unanswered tool call is NOT the same failure -- it's the completely normal shape of a mid-conversation turn seconds after it lands, before the agent runtime sends the next request with the tool's result. Flagging that immediately would false-positive on every healthy in-flight tool-calling conversation. Adds a "stalled" badge, server-computed (Product Doctrine: record the fact where it's known, not multi-signal inference in the client): a conversation whose last turn didn't end in a clean "stop" (either a genuinely incomplete/truncated stream, or a tool call still unanswered) AND 5+ minutes have passed since lastSeenAt with nothing having continued. Never true while isActive -- an actively streaming/polling conversation is by definition not abandoned. resolveTurnCompletionState (packages/db/responsesContinuationStore.ts) is a pure, permanently-cached function of one call-log's own already-persisted artifact -- same immutability argument and cache shape as the existing sibling isGenuineContinuationTurn (an artifact never changes once detailState flips to "ready"). resolveConversationStalledState combines that with the live isActive/lastSeenAt facts (never cached, since those change over time) to decide the badge. Wired through the existing annotateConversationRow shared by both the list and single-conversation routes -- no new API surface, no schema change. Tests: tests/unit/responses-continuation-store.test.ts -- 10 new cases covering all three completion states (stop / tool_call_pending / incomplete, including the exact truncated-collector shape from the live incident), the 5-minute grace-period boundary in both directions, isActive always overriding to false, and a clean stop never flagging regardless of elapsed time. Full file: 23/23 passing. Verified live against the real flagged conversation from the incident: /api/conversations now returns isStalled: true for it. eslint / tsc --pretty false -p tsconfig.typecheck-core.json clean on all touched files.
hartmark
added a commit
to hartmark/OmniRoute
that referenced
this pull request
Sep 4, 2026
…at never reached a clean stop) into dev/omniroute-dev-combined
hartmark
added a commit
to hartmark/OmniRoute
that referenced
this pull request
Sep 4, 2026
…at never reached a clean stop) into dev/omniroute-dev-combined
hartmark
added a commit
to hartmark/OmniRoute
that referenced
this pull request
Sep 4, 2026
…at never reached a clean stop) into dev/omniroute-dev-combined
hartmark
added a commit
to hartmark/OmniRoute
that referenced
this pull request
Sep 4, 2026
…at never reached a clean stop) into dev/omniroute-dev-combined
hartmark
added a commit
to hartmark/OmniRoute
that referenced
this pull request
Sep 4, 2026
…at never reached a clean stop) into dev/omniroute-dev-combined
hartmark
added a commit
to hartmark/OmniRoute
that referenced
this pull request
Sep 5, 2026
…at never reached a clean stop) into dev/omniroute-dev-combined
…onnections (diegosouzapw#12470) Merged. Clean, surgical fix with its own regression guard. `ensureProviderConnectionsColumns()` reconciles the base columns that later data migrations assume, but `last_ping_at` / `last_pinged_reset_key` were only ever created by `123_quota_auto_ping` — so a lineage that skipped it kept a table that the quota auto-ping writes cannot target. Adding them to the reconciliation list is exactly the right place. Validated on `release/v3.8.51`: `tests/unit/db-schema-columns-split.test.ts` 10/10, including your new `back-fills last_ping columns on a pre-123 lineage` case and the idempotency re-run. `typecheck:core` clean, `check-file-size` OK. The `changelog.d/fixes/` fragment was already correct. Thank you — this is the shape a fix should have: root cause named, minimal diff, test that fails without it.
… object (diegosouzapw#12494) (diegosouzapw#12771) Merged, with one adjustment. Confirmed the bug end to end: `PricingTab.tsx:369` types the payload as `{ error?: string }` and feeds it to `new Error(errorPayload.error || ...)`, so the `{ message, details }` object landed in the toast as `[object Object]` — exactly what diegosouzapw#12494 reported. The one change I made before merging: `validation.error.message` is the fixed constant `"Invalid request"` (see `validateBody` in `src/shared/validation/helpers.ts:44`), so it would have swapped an unreadable toast for an uninformative one. The repo already has `formatValidationMessage()`, added in diegosouzapw#10849 for precisely this case — it returns `"field: reason"` naming the first offending field. Merged with that instead, so a bad pricing value now says which field it was. Validated on `release/v3.8.51`: `typecheck:core` clean, `check-file-size` OK. Rebased onto the release branch — the PR was cut from `main`, which is ~3695 commits behind the active branch. Thank you for the report and the fix.
…FILES description (diegosouzapw#12505) (diegosouzapw#12769) Merged, with the fix moved to where the bug actually lives — and thank you, because the issue analysis in diegosouzapw#12505 is what made that possible. The diagnosis was right: `FeatureFlagsGrid.tsx:422-428` renders descriptions through a plain `t()`, so next-intl compiles the value as ICU and a bare `<name>` parses as an unknown rich-text tag. But the branch changed `src/shared/constants/featureFlagDefinitions.ts` — the TS default, which is the `flag.description` **fallback** rendered raw, never through ICU. Two consequences: the reported bug stayed live (all 42 locale files still carried the raw tag — `grep -l "profiles/<name>/settings.json" src/i18n/messages/*.json` returned 42, and 0 for the escaped form), and the quotes would have shown up literally in the one place that string does render. So this merge reverts the TS default to the raw path and applies the ICU escape to the 42 locale files instead — follow-up 1 from your issue, inverted to hit the file that matters. I also added follow-up 2 as a real guard: `tests/unit/feature-flag-description-icu-parse-12505.test.ts` compiles every `featureFlags.definitions.*` message in every locale through `intl-messageformat` (the parser next-intl uses) and asserts the placeholder renders as a literal `<name>`. Verified red-then-green — reverting `en.json` alone fails both cases; restored, 2/2 pass. Validated on `release/v3.8.51`: all locale files re-parse as valid JSON, `typecheck:core` clean, `check-file-size` OK. `i18n:check` drift is pre-existing on the tip, unrelated. Closes diegosouzapw#12505.
…diegosouzapw#12575) (diegosouzapw#12764) Merged, with one sentence removed. The `socket.yml` half checks out: the file exists at the repo root, is `version: 2`, and its `projectIgnorePaths` really do list `tests/`, `_tasks/`, `_references/`, `_ideia/`, `_mono_repo/`, `docs/` — so the paragraph describes the config accurately. The closing sentence did not: there is no `.github/workflows/socket-dev.yml` in this repo (`ls .github/workflows | grep -i socket` is empty), and nothing auto-opens `supply-chain-review/` issues. Per the documentation-accuracy rule in `AGENTS.md` — every path and workflow named in docs has to survive an `rg`/`ls` — I replaced it with what is actually true: the scan is driven by the Socket GitHub App reading `socket.yml`, not by a workflow here. Everything else merged as written. Thanks — pointing readers of SECURITY.md at the scanner config was a real gap.
…zapw#12570) (diegosouzapw#12706) Merged, with the threat model kept and two unverifiable claims dropped. The core warning is correct and worth having in both files: `/var/run/docker.sock` is a host-root trust boundary, the `cli` profile must not be published beyond `127.0.0.1`, and no extra host mounts belong in it. That is now in `docker-compose.yml` next to the mount and in the DOCKER_GUIDE. Two things I changed before merging, both `AGENTS.md` documentation-accuracy calls: 1. **The stated purpose.** The socket is not mounted so OmniRoute can "launch short-lived codex/claude-code/droid/openclaw containers" — I could not find any container-spawn path. It is there for the in-container auto-updater: `src/lib/system/autoUpdate.ts:236` probes for `/var/run/docker.sock` and skips the Docker path when it is absent, and the mount sits right beside `AUTO_UPDATE_HOST_REPO_DIR`. Rewrote the sentence around that and cited the file. 2. **Item 3, the audit log.** "recorded in the server log with the called tool, the prompt digest (not content), and the spawned image SHA" — no such logging exists (`grep -rn "prompt digest\|promptDigest\|imageSha" src/ open-sse/` is empty). A security doc promising forensics that are not implemented is worse than one that stays quiet, so I removed the item rather than soften it. The `MITM-TPROXY-DECRYPT.md` and `SUPPLY_CHAIN.md` cross-references both resolve and stayed. Thanks — the docker.sock boundary genuinely was undocumented.
) (diegosouzapw#12667) Merged, with the markdown repaired. Checked every row against `src/lib/gamification/xp.ts:138` — the table now matches `XP_REWARDS` exactly, keys and values, and the descriptions are the JSDoc lines verbatim. The old table was documenting actions that do not exist (`badge_earned`, `streak_milestone`, `referral`, `model_diversity`, `compression_use`, `skill_use`) and missing the three that do (`model_switch`, `invite_redeem`, `streak_bonus`). Good catch. Two formatting fixes before merge: the action names were padded inside the code spans (`` `request ` ``), which renders the trailing spaces as part of the identifier; and the unrelated MCP-tools table below had its header row flattened, losing the column alignment. Restored both and ran Prettier — the file is clean now. Thank you for reconciling this against the source instead of guessing.
…iegosouzapw#12703) Merged. Verified all four pins resolve on npm before landing: ``` @openai/codex@0.153.2 0.153.2 @anthropic-ai/claude-code@2.1.260 2.1.260 droid@0.212.0 0.212.0 openclaw@2026.9.1 2026.9.1 ``` The reproducibility argument holds — a floating `@latest` in a cached Docker layer means two builds of the same commit can ship different toolchains, and that is exactly the class of drift that makes a CI failure unattributable. Worth flagging for whoever maintains this next: pinning trades drift for staleness, so these four now need a periodic bump or the image ships increasingly old CLIs. The comment block you added explains the why, which makes that bump a safe mechanical change instead of a judgment call. Rebased onto `release/v3.8.51` (the PR was cut from `main`, ~3695 commits behind). Thanks.
Merged. One line in `overrides`, low blast radius, and pinning a transitive that every build tool reads is defensible on its own. Validated on `release/v3.8.51`: `package.json` re-parses, `typecheck:core` clean, `check-file-size` OK. For future dependency pins, a line in the body about what the floating range actually broke (a specific build failure, a CVE, a resolution conflict) makes these reviewable without guessing. Thanks.
… to consumers (diegosouzapw#11544) (diegosouzapw#12699) Merged — one line, zero risk, and it costs nothing to have. One caveat recorded so nobody later reads this as "diegosouzapw#11544 is solved": npm resolves config from the *installing* project's directory, the user config and the global config — it does not read the `.npmrc` shipped inside a dependency's tarball. So `legacy-peer-deps=true` traveling in the package will not change how `npm install -g omniroute` resolves peers on the consumer side. Our own `scripts/build/postinstall.mjs` does shell out to `npm rebuild` / `npm install better-sqlite3`, but with cwd set to `dist/`, so the package-root `.npmrc` is not in scope there either. Keeping it anyway: it makes the published tree self-documenting, and someone debugging inside an extracted package gets the same retry budget we use in CI. But diegosouzapw#11544 (`npm install -g omniroute` failing on Windows, "root cause unclear from log") still needs the actual `npm-debug.log` from the reporter before it can be closed. Rebased onto `release/v3.8.51`; `package.json` re-parses and the `files` array kept both `config/i18n.json` and the new entry. Thanks.
…cuting gate test (diegosouzapw#12581) (diegosouzapw#12652) Merged as a reduced diff, and worth recording why. The two-arm `ALWAYS_PROTECTED` read this PR proposed had already landed in diegosouzapw#12605 while this branch was open — the tip carries `coveredByAlwaysProtected()` with both arms and the new error wording. I ran the gate on the current tip to be sure: `PASS — all security tier annotations match routeGuard.ts`. Merging the whole branch would have reintroduced the same logic under a different comment. What was genuinely missing, and is what merged: - **`tests/unit/openapi-security-tiers-gate.test.ts`** — executes the real gate and asserts exit 0 with no "NOT covered" line. diegosouzapw#12605 fixed the defect but left no guard, so the LOCAL_ONLY-arm bug (diegosouzapw#12350) could reappear on the ALWAYS_PROTECTED arm exactly as it did the first time. 1/1 green. - **The parse guard** — `ALWAYS_PROTECTED_PATTERNS.length === 0` now fails the constant-parse check with its own count in the message. Without it, a regex array that stops parsing degrades into "every pattern-covered route is an annotation mismatch" instead of saying so. A note for the record: my first read of this PR was wrong. I ran the gate in the main checkout, which was 11 commits behind `origin/release/v3.8.51`, saw the pre-diegosouzapw#12605 failure, and classified this as fixing a live red. It was not — the checkout was stale. Corrected before anything was merged.
…ogs (diegosouzapw#12150 P2 item 7) (diegosouzapw#12710) Merged. Focused, correct, and tested. `recordRejectedRequestUsage` runs on the path where the request never reached the guardrail chain — circuit-breaker-open and combo-exhausted rejections — so the video-bridge guardrail never got the chance to rewrite the transcript, and the raw cues went straight into `call_logs`. Routing the body through `redactVideoTranscriptFieldsForLog` at the persistence boundary is the right place: a no-op clone for non-video bodies, structured field substitution for video ones, and not bypassable by cue content. The `requestBody == null ? requestBody : …` guard keeps the existing "no body available" case behaving exactly as before, which the neighbouring test still covers. Validated on `release/v3.8.51`: `tests/unit/rejected-request-usage.test.ts` green, including the new case asserting the secret cue text does not survive into the persisted detail and that the field reads `[redacted-video-transcript]`. `typecheck:core` and `lint` clean.
…l store (base-red diegosouzapw#12581) (diegosouzapw#12671) Merged. This removes the cause that diegosouzapw#12607 had to freeze. `react-hooks/set-state-in-effect` on this file was living in `config/quality/eslint-suppressions.json` as a frozen count of 1 — the lint was green because the violation was suppressed, not because it was gone. `useSyncExternalStore` is the sanctioned shape for exactly this problem: `getServerSnapshot` supplies the SSR-safe default, `getSnapshot` reads localStorage after hydration, and the tree commits once instead of twice. The `storage` listener keeping other tabs in sync is a real bonus. The detail that makes this correct rather than merely lint-clean: you kept "hide for now" and "hide forever" as separate concepts — `usageGuideHiddenForNow` stays per-mount local state while only the persisted dismissal goes through the store. A naive conversion would have collapsed them and made the temporary hide survive a reload. Three things I added before merging: 1. **Dropped the `react-hooks/set-state-in-effect` entry from the suppressions file.** With the cause gone it becomes a stale allowlist entry, which is what the Fase 6A.3 stale-enforcement is built to flag. Verified: `eslint` on the file now reports only the 6 pre-existing `no-unused-vars`, which stay frozen. 2. **Updated the rationale comment above the hook** — it still described "correct it client-only, after hydration, in an effect", which is the shape you just removed. 3. **Rebaselined `combos/page.tsx` 5018 → 5066** in `file-size-baseline.json` with a dated annotation. The +48 lines are the module-scope store helpers; the cap is pre-authorized for legitimate growth and this is as legitimate as it gets. Validated on `release/v3.8.51`: `check-file-size` OK, `lint` clean, `typecheck:core` and `check:dashboard-typecheck` clean (207 pre-existing, all within baseline).
…video turns (diegosouzapw#12150 P2b) (diegosouzapw#12707) Merged, with one column-reconciliation gap closed. The fail-closed reasoning is right and the comments carry it well: a stored snapshot whose cues were replaced by `[redacted-video-transcript]` must not be rehydrated as continuation history, because forwarding placeholder text upstream as if it were the client's real turn is worse than making the client resend. Treating it exactly like `previous_response_not_found` means no new client-visible behaviour to document. Migration 173 does not collide — the tip runs to 172. **What I added:** `video_content_removed` to `ensureCallLogsColumns` in `src/lib/db/schemaColumns.ts`, plus a case in `tests/unit/db-schema-columns-split.test.ts`. `resolvePreviousResponseState` now SELECTs that column on every `previous_response_id` lookup. Migration 173 creates it, but this repo carries a separate reconciliation path for lineages that skipped a migration — and on such a database the SELECT would throw `no such column: video_content_removed` instead of failing closed. That is the same hole diegosouzapw#12470 closed for `provider_connections.last_ping_at` earlier today, so the pattern was fresh. Verified red-then-green: stubbing the new reconciliation out drops the suite to 8/9; restored, 9/9. Validated on `release/v3.8.51`: `responses-continuation-store`, `save-call-log-persistence`, `video-bridge-log-redaction` and `db-schema-columns-split` all green (54 focused tests, 0 failures). `typecheck:core` and `lint` clean. The integration run logs `[DB] Added call_logs.video_content_removed column`, which is the reconciliation firing on a fresh test database.
…suite (base-red diegosouzapw#12581) (diegosouzapw#12670) Merged. It does what it says, and it also uncovered something — details below so the follow-up is not mistaken for a regression from this PR. Measured on `release/v3.8.51`, `tests/integration/chat-pipeline.test.ts`: | | line 580 | line 994 | line 1599 | |---|---|---|---| | tip | `410 !== 200` | `410 !== 200` | `502 !== 200` | | tip + this PR | passes | passes | passes | All three were retired model ids reaching the router and coming back 410/502. Swapping them for live ones is exactly the right fix and takes the suite from 25/28 to 27/28. **The one that remains, and why it is not yours:** with the 410 gone, `chat pipeline persists Codex responses cache and reasoning tokens to call logs` now runs past `assert.equal(response.status, 200)` and reaches line 592, where `callLog.provider` is `openai` and the test expects `codex`. That assertion was simply never reached before — the 410 short-circuited the test at line 580. I checked whether the model id chosen here was the cause, since `gpt-5.6-sol` is declared by 12 providers (`openai`, `github`, `cursor`, `kiro`, …). It is not: re-running with `gpt-5.3-codex-spark`, which only the `codex` provider declares, produces the identical `openai !== codex`. So it is provider resolution or the `seedConnection("codex")` harness, not catalog ambiguity. I reverted that experiment — this merged exactly as you wrote it. Filing that as its own issue with the trace.
…iegosouzapw#12834) Native Codex passthrough (POST /v1/responses, provider=codex) was unconditionally excluded from prompt compression, writing only skip_reason='excluded' analytics rows. Prompt compression now depends only on the operator exclusions list; reactive compaction and combo overflow fail-fast intentionally still bypass (prompt-only scope). Closes diegosouzapw#12793 Regression guard: tests/unit/codex-prompt-compression-passthrough.test.ts
hartmark
added a commit
to hartmark/OmniRoute
that referenced
this pull request
Sep 6, 2026
…at never reached a clean stop) into dev/omniroute-dev-combined
…iegosouzapw#12870) opencode v2 loads plugins through a contract the existing @omniroute/opencode-plugin cannot satisfy: v1 exports plugin factories with an auth/provider/config/tool hook object, v2 expects a default define({id, setup}) carrying catalog and integration domains. One package would have to satisfy both loaders from a single entrypoint. An opencode v2 install therefore has no route to an OmniRoute gateway at all: no model discovery, no combos, no enrichment. This adds @omniroute/opencode-plugin-v2, a self-contained package. The v1 plugin is untouched, so v1 users see no move, no migration and no breaking version. The two packages deliberately share no code and no release: the mapping logic here began as a port of v1's and now lives in this package, which keeps either one free to change without a coordinated publish. The plugin publishes models, combos and auto-combos into the host catalog, refreshes them lazily behind a 300s TTL, and keeps serving the last known catalog from an on-disk snapshot when the gateway is unreachable. Publishing is staged: models and combos are what a catalog is, so they go out as soon as they are known, while auto-combos, the provider list and the enrichment overlay fold into the snapshot when they land. Gating the publish on all of them made the catalog hostage to the slowest source — a gateway that accepts the connection and never answers /api/combos/auto left everything unpublished until that fetch timed out, which is longer than a short-lived host stays alive. Display names carry what the gateway knows about a model: the upstream provider it routes to, whether it is free, and the budget that comes with it. Those parts were already fetched and then dropped, so two connections selling the same model looked identical in the picker. The provider prefix can be turned off with `providerTag: false`. The on-disk snapshot carries that overlay too, under a size cap, so a cold start opens on named models rather than raw ids. The host is asked to reload only when the catalog or the overlay actually moved, never once per refresh window. The gateway key comes from the host credential store when one is connected, so connecting the integration from opencode is enough and no secret needs to sit in opencode.json; a plugin option and an environment variable remain as fallbacks, and a host too old to expose a credential store still loads. Nothing is silent when a key is missing or refused: an absent key is named once at startup with the three ways to supply one, and an enrichment source the gateway rejects is reported per endpoint with what the catalog loses. Those three failures used to be empty catch blocks, which turned a management token the gateway refuses into a catalog of raw model ids with no explanation. Tool calling to Gemini keeps working. Gemini answers 400 INVALID_ARGUMENT for an entire request whose tool declarations carry $schema, $ref or additionalProperties. The v1 plugin handled it by wrapping fetch and rewriting the JSON body; v2 does it on the language model, where the tools are still structured data, and only for Gemini models of this provider. It can be turned off with geminiSanitization: false, and a host exposing no aisdk domain loads without it. The catalog contract itself is a moving target, so the plugin adapts to the host instead of assuming one shape. The released CLI keeps the aisdk package, the endpoint (as settings.baseURL), the request headers and the variant options directly on the model and provider; the current SDK types keep the same information inside an api block. Writing only the api block yields a catalog the released CLI lists but cannot route. Rather than key off a version list that goes stale on the next release, the plugin reads the shape the host seeds into the catalog draft and publishes accordingly: a seed with a top-level package and no api block gets both field sets, a seed with an api block gets that block alone, and an undisclosed seed gets both. None of the legacy keys collide with a key of the current types, so the two shapes coexist on one object, variants included. Four v1 behaviours are deliberately not carried over, because v2 either owns them or no longer needs them: the plugin-side debug log (the host has its own logging), the compression-metadata suffix on combo names, the MCP auto-emit (the v2 host owns MCP), and the omni-sync command plus its background timer (the TTL and a content fingerprint drive catalog.reload instead). A refresh never downgrades what is already published: the previous overlay is carried forward until the new one lands, so names, pricing and the usable filter no longer drop out for the length of every TTL window. The disk snapshot is read after the credential is resolved, because it is keyed by that credential — reading it earlier looked up the identity the options carry rather than the one in use, and rejected a perfectly good catalog exactly when the gateway was down. The tool-schema cleaner now knows where a schema ends and a property name begins. Stripping keywords by name anywhere in the tree deleted a tool parameter called `ref` while leaving it in `required`, handing the model a schema it could not satisfy; a `$ref` it cannot resolve now forwards the tool untouched instead of widening it to accept anything. Gemini detection is anchored on the model family, so `gemini-compatible-proxy` is no longer treated as a Gemini model. A source the gateway refuses is reported on the library entry point as well, not only through the plugin, so the usable-provider filter can no longer disable itself in silence. `providerId` is bounded to a safe character set because it reaches a filesystem path, `hiddenModels` covers combos as it already covered models, the Anthropic block gets the gateway root rather than a doubled `/v1`, an unparseable tool schema forwards the tool instead of failing the request, and the package typechecks under the same settings as the v1 plugin. CI mirrors the existing plugin workflow: install, build and test on Node 22 and 24, for both packages. The plugin SDK stays pinned, and the host-shape assertions carry the risk of a contract move rather than a check against a rolling upstream tag. Co-authored-by: Max <maxmad64@gmail.com>
…ibution (diegosouzapw#12772) * chore(ci): guard commit identity in pre-commit to stop author misattribution Two windows of commits in this checkout were signed with the wrong identity, both caused by an identity override left behind by an automated session: 2026-08-13..26 (name "Xiangzhe" + @backryun's e-mail, 237 commits) and 2026-08-29..09-02 (name "Markus Hartung" + the maintainer's e-mail, 59 commits). The .mailmap repairs the record after the fact; this gate stops the next window. The gate is opt-in per machine via omniroute.expectedName / expectedEmail — with no config it exits 0, so contributors who clone the repo are never affected. It blocks three things: a committer that is not this machine's identity (which is what BOTH windows looked like — in August neither the name nor the e-mail was the maintainer's, so checking only their e-mail would have missed it), an author carrying the maintainer's e-mail under someone else's name, and any address listed in omniroute.legacyEmail. Crediting a contributor with `git commit --author="Name <their@email>"` keeps working, since the rule targets the committer and the maintainer's own address. * test(ci): isolate the identity gate's test from the ambient git config The "stays inert when the machine has not opted in" case read the real global config, so on a machine that HAS opted in (omniroute.expectedEmail set — the maintainer's own boxes, where this gate matters most) the gate correctly refused a synthetic contributor identity and the test failed. It only passed on a clean CI runner. Neutralising GIT_CONFIG_GLOBAL/GIT_CONFIG_SYSTEM makes the opt-in state come solely from what the test injects, so the suite is deterministic on both an opted-in and a clean machine.
…2770) Validado numa worktree combinada com as 16 PRs desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-file-size e check-changelog-integrity OK, complexity 2788/3218 e cognitive 1261/1437 (ambos sob a baseline), ESLint 0 erros nos 152 arquivos alterados, 771 testes unitários focados, 49 de integração e a suíte vitest:ui completa (2149) verdes. O `65536` era o 16º posicional de um helper com 15 parâmetros — `TS2554` vivo no tip (`open-sse/executors/glm.ts:244`, confirmado aqui antes do board). O teste de guarda de aridade é o que impede a reincidência: ele checa a assinatura do helper e o call site, não o comportamento, que é exatamente onde o erro morava. Obrigado por isolar isso do diegosouzapw#12711 em vez de deixar o `glm.ts` viajar junto com pin/combo-split/moonshot.
…agment (diegosouzapw#13200) Um caractere. O fragmento do diegosouzapw#13097 subiu sem o `- ` inicial e derrubou o `Merge integrity` para todo mundo que veio depois. Terceira ocorrência da mesma causa nesta release; a anterior foi o `reset-aware-model-family.md`, que o diegosouzapw#12711 consertou de carona.
Contagem de migrations 171 → 172 após a `175_call_logs_provider_stats_indexes.sql` do diegosouzapw#12832. Medido com `ls src/lib/db/migrations/*.sql | wc -l`. Falha minha de processo: depois da onda 1 desta leva eu medi file-size, api-typecheck, changelog-integrity e colisão de migration — não o `check:docs-counts`. O drift ficou vivo até a onda 2 esbarrar nele. 41 mirrors de `llm.txt` regenerados pelo script do projeto. Aprovado por você para tocar `AGENTS.md`, mesma classe do diegosouzapw#12970.
…thout a documented free tier (diegosouzapw#12744) Validado numa worktree combinada com a onda de roteamento/free-tier desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (288), check-file-size OK após rebaseline, 84/85 nos testes focados no runner Node. Parar de honrar `isFree:true`, `:free` e `0/0` vindos do upstream **antes** de consultar o catálogo é a inversão certa: hoje um provider sem tier livre documentado consegue se declarar grátis e o listing diverge do roteador `auto/*`. Checar as heurísticas depois do hit de catálogo fecha a porta sem quebrar o caminho de linhas custom locais, que continuam confiáveis pelo caminho próprio. O `isFreeModel("or", …)` com um alias que não existe é o tipo de bug que passa despercebido porque falha silenciosamente para o lado permissivo. **Integração:** o `decideHidePaid` que o diegosouzapw#12795 extraiu passou a usar o seu `isFreeForProvider` por id, em vez de OR-ear os dois aliases num único `freeProvider`. A forma por id é a garantia que esta PR estabelece, então ela prevaleceu.
…nifest (diegosouzapw#12786) Validado numa worktree combinada com a onda de roteamento/free-tier desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (288), check-file-size OK após rebaseline, 84/85 nos testes focados no runner Node. Metadado de display derivado da mesma fonte da decisão, com um gate STRICT que quebra o CI se a contagem do manifesto divergir do catálogo — é o detalhe que impede a tag de virar mentira daqui a três meses. Registrar que 77 entradas de catálogo viram 76 no manifesto porque o `arcee-ai` ainda não tem entrada no registry é exatamente o tipo de discrepância que costuma virar bug fantasma.
…ty/breaker (diegosouzapw#12794) Validado numa worktree combinada com a onda de roteamento/free-tier desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (288), check-file-size OK após rebaseline, 84/85 nos testes focados. Tratar breaker aberto como equivalente a fechado no scoring de snapshot é pior que não pontuar: afirma saúde onde há falha conhecida. Trocar as três constantes neutras por valor observado é a correção, e manter preço e orçamento fora do escopo mantém a PR revisável. **Integração:** o `computeSnapshotWeights` conflitou com o diegosouzapw#12731, que adiciona peso de `reliability` a partir de `failureRate`/`errorRate`. Os dois cobrem chaves diferentes e compõem — ficaram ambos: reliability do diegosouzapw#12731, health via breaker e quality desta PR, quota neutro nos dois. Nenhum dos dois lados foi descartado.
…diegosouzapw#12731) Validado numa worktree combinada com a onda de roteamento/free-tier desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (288), check-file-size OK após rebaseline, 84/85 nos testes focados. `reliability-first` que não pesava reliability é o defeito mais constrangedor possível num mode pack, e a causa é clara: `modePacks.ts:13` substituía os defaults por inteiro. Financiar os novos pesos com `quota`/`costInv`/`tierPriority` mantendo cada pack somando 1.0 é a parte que exige cuidado e você fez. Manter `quality-first` em 0.03, igual ao default, para que ele não fique mais fraco que `balanced`, é o tipo de detalhe que só aparece quando se checa a tabela inteira. **Integração:** conflitou com o diegosouzapw#12794 no `computeSnapshotWeights`; os dois compõem e ambos ficaram.
…iegosouzapw#12790) Validado numa worktree combinada com a onda de roteamento/free-tier desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (288), check-file-size OK após rebaseline, 84/85 nos testes focados. Duas cópias da mesma ordem de provider com um "keep in sync" implícito é dívida que cobra juros a cada provider novo. Uma definição com re-export nos dois lados resolve a classe. O `xao/*` ordenando depois de todo provider conhecido em vez de junto do `xai-oauth` é um sintoma concreto de que a duplicação já estava divergindo. **Integração:** `scripts/quality/run-all-gates.mjs` conflitou com o `check:pricing-freshness` que entrou pelo diegosouzapw#12792 na mesma onda. Aditivo — os dois gates coexistem.
…s, pooled latency bootstrap, fresh tier cache (diegosouzapw#12792) Um modelo grátis fora da tabela herdando $5/$15 por milhão e afundando no roteamento cost-aware é o defeito mais caro desta onda: silencioso, e inverte exatamente a decisão que o operador quer. Parar de chutar 1500ms de latência para modelo desconhecido e usar a mediana observada do pool — com contador de quantas vezes o chute dispara — é trocar heurística por medição do jeito certo. O contador é o que permite saber se valeu. Revalidei após reconstruir a branch sobre o tip: **33/33** nas suítes da PR, typecheck:core limpo, `check-api-typecheck` OK (289). **Duas integrações:** 1. `computeSnapshotWeights` conflitou com o diegosouzapw#12794 (health via breaker + quality), já mergeado. Os dois compõem e ambos ficaram: o seu termo de `reliability` — que era a única chave que o caminho de snapshot ainda ignorava — mais o health observado e o quality do diegosouzapw#12794. 2. `scripts/quality/run-all-gates.mjs` conflitou com o `check:provider-order-sync` do diegosouzapw#12790. Aditivo, os dois gates coexistem. **Nota de dívida:** o `virtualFactory.ts` cruzou o teto de 1200 linhas pela primeira vez (1187 → 1207) somando esta onda. Congelei em vez de dividir e registrei os dois candidatos a extração na justificativa — `computeSnapshotWeights` (~85 linhas) e o grupo de elegibilidade de credencial (~70). Qualquer um dos dois volta o arquivo para baixo do cap.
…op reasons (diegosouzapw#12795) Um pool `auto/*` vazio que não diz por que está vazio é a pior forma de falha: o operador vê ausência e não sabe se é config, cota ou catálogo. Registrar qual estágio removeu quantos, e carregar `dropReason` em cada entrada de `quotaHealth.providers`, transforma silêncio em diagnóstico. Os três defeitos achados de carona valem tanto quanto a feature — em especial a cota livre recorrente sem teto sendo tratada como desconhecida em vez de segura, que é justamente o caso que mais aparece. Revalidei após reconstruir sobre o tip: **15/15**, typecheck:core limpo, `check-api-typecheck` OK (289). **Uma mudança minha no `catalogPaidFilter.ts`.** Mantive a sua extração — ela é mais limpa que o predicado inline — mas troquei o corpo para usar `isFreeForProvider` por id. O módulo OR-eava `providerHasFreeModels(resolved) || providerHasFreeModels(canonical)` num único `freeProvider` e depois testava `isFreeModel` em cada um. Com isso, um id cujo próprio provider não documenta tier livre passa a ler como grátis sempre que o alias irmão documenta — que é exatamente o buraco que o diegosouzapw#12744 fechou e já está no tip. Por id preserva a garantia; nenhuma outra linha do módulo mudou. **Dívida registrada:** o `virtualFactory.ts` foi de 1187 para 1219 somando esta onda e cruzou o teto de 1200 pela primeira vez. Congelado com os candidatos a extração nomeados na justificativa.
…diegosouzapw#13044) Batch 1 of the locale expansion: Greek, Croatian, Serbian, Lithuanian, Estonian, Latvian, Slovenian, Maltese and Irish across the dashboard catalog, docs mirrors, CLI catalog, README, locale index and the site. 42 → 51 locales. Also fixes the ICU literal escape the translation backend dropped around angle placeholders, four translations that invented or renamed a placeholder, the language bars that linked to mirrors that do not exist, and the migration count drift (171 → 172).⚠️ base-red inherited: diegosouzapw#12732 — the four unit shards and Fast Quality Gates fail identically on unrelated PRs cut from the same base.
…e History tab (2.9) (diegosouzapw#12677) * feat(dashboard): pure model to compare two orchestration runs * feat(dashboard): compare-mode selection in the History grid * feat(dashboard): side-by-side comparison panel in the History tab (2.9) * chore(dashboard): compare-runs i18n + changelog * fix(dashboard): compare-panel loading state, height bound, delta legend, ARIA level Final-review fix wave for PR-A (Orchestration Canvas Fase 3): - Give each compare-panel side an explicit fetch status (loading/ok/error) so the Events metrics row shows "—" instead of a misleading real "0"/delta while a side is still loading or after its fetch failed. - Bound the compare panel's height (max-h-[45vh], overflow-y-auto, shrink-0) so a run with many activities can no longer collapse the History grid to zero height. - Clear the compare selection when the History preset changes, since a stale pick can fall outside the new range. - Add a delta-column legend (new compareDeltaLegend i18n key, translated into all 41 non-English locales) so operators know which side a positive delta favors. - Use role="status" (not role="alert") for the informational compareDifferentIdentity banner, reserving role="alert" for actual per-side fetch failures. - Restore ro.json's compareCost to the true cognate "Cost" (was distorted to "Cheltuieli" to dodge a byte-identical-to-English heuristic); audited the other 40 locales for the same pattern across the 9 compare* keys, no other instance found. * refactor(dashboard): split compare-runs/history-tab functions to clear complexity ratchets CompareRunsPanel (92 lines, max-lines-per-function) and HistoryTab (85 lines, same rule) exceeded the 80-line function cap; compareRuns.ts's a2aEventsFrom exceeded the cognitive-complexity cap (16 > 15). Extract pure/presentational helpers (ComparePanelHeaderBar, SideErrorRow, ComparisonMetrics, a2aEventFrom, HistoryStatusRows, refreshNowMsOnActionDone) with no behavior, DOM, i18n or aria change. --------- Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
…zapw#12392) (diegosouzapw#12983) * fix(dashboard): keep the first failure timestamp in sourceStale buildSourceStatuses stamped nowIso on every failing source at every poll, so the stale indicator reported "since the last poll" instead of the first failure — and, because snapshotContentKey serializes sources, the snapshot identity churned on every tick while any source was down. The failing branches now reuse the staleSince already held by that source in the previous status list, via the functional setStatuses updater (no ref read during render, no setState inside an effect body). Refs diegosouzapw#12392 * fix(dashboard): flag a source that starts failing after it had data buildRootAndSourceEdges only materialized a placeholder SourceNode when the failing source had no node at all. A source that already had work nodes and then started failing (or went offline) kept its healthy-looking SourceNode forever: no ⚠, no stale styling, no `sourceStale` line — the operator saw a normal source while it was actually broken. Now every non-ok/offline source is flagged: when its SourceNode is missing the placeholder is created as before; when it exists, the node is replaced by a copy carrying `sourceIssue` and `staleSince`. The copy (never a mutation) keeps the function pure — the original object is still referenced by the caller's `parts`, the same trap the droppedByState aliasing fix covered. Tests: three cases in tests/unit/ui/orchestrationModel.test.ts — existing node starting to fail (flags set, work nodes kept, no duplicate node, input object untouched), existing node going offline (no invented staleSince), and a healthy source staying free of both fields. Refs diegosouzapw#12392 * fix(dashboard): canvas polish batch (diegosouzapw#12392) Seven pointwise fixes on the Orchestration Canvas, each covered by a test: 1. Debounce x chip race: every chip/clear write in OrchestrationToolbar now cancels the pending search timer first. Left armed, it fired ~300ms later with a setParams closed over the pre-chip query string and silently reverted the chip. 2. The search input carries an aria-label (searchPlaceholder) — the placeholder alone is not an accessible name. 3. parseCsvSet trims each token, so `?state=running, failed` parses like the unpadded form instead of dropping the padded value. 4. toggleCsv was duplicated in the toolbar and the page client; both now import the single definition from the new model/urlParams.ts (pure, never mutates its inputs). 5. AgentsTab tells "nothing running" apart from "the filter matched nothing": with an active filter and no work node it renders noMatches + a clear-filters button instead of the setup CTAs, which would be wrong advice there. 6. Particle cap: orchestrationToFlow stamps `particles` on every edge and turns it off above PARTICLE_EDGE_CAP (40) simultaneously active edges — StatusEdge then renders the colored stroke without its 3 SMIL particles per edge. 7. The drawer's error banner clears when an action succeeds, so a recovered failure does not stay on screen. Only `noMatches` is added to en.json here; the other locales are task B4. Refs diegosouzapw#12392 * chore(dashboard): canvas polish i18n + changelog Real translations for orchestration.noMatches in the 41 non-English locales, each one written against that file's own neighbouring keys (emptyTitle, stateRunning, searchPlaceholder) so the wording for "task" and "filter" matches what the locale already uses. No i18n:sync-ui, no __MISSING__ left. Adds the changelog fragment for the nine PR-B fixes. Closes diegosouzapw#12392
…ck (diegosouzapw#12880) Validado numa worktree combinada com a onda de streaming desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 88/88 nos testes focados. `agent_message` chegando num fallback de Chat Completions é um item que o cliente não sabe interpretar; mapear ou descartar é a escolha certa, e escolher por item em vez de derrubar a resposta inteira mantém o fallback útil.
…ream stays silent (diegosouzapw#12828) Validado numa worktree combinada com a onda de streaming desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 88/88 nos testes focados. O diegosouzapw#12151 cobriu só metade: passthrough emitia o chunk final de usage, translate calculava a estimativa **depois** de fechar o stream, então o número só chegava ao log do servidor e nunca ao cliente. Fechar essa metade é o que faz a feature existir de fato. Não emitir segundo chunk quando o upstream já mandou usage real é o detalhe que impede a correção de virar contagem dobrada.
…p-dropped event array (diegosouzapw#12718) Validado numa worktree combinada com a onda de streaming desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 88/88 nos testes focados. Reconstruir o resumo a partir de um array que o próprio coletor já truncou por cap produz um resumo que parece completo e não é — pior que resumo ausente, porque não se distingue. Parar de reconstruir dali é a correção.
…ycle/heartbeat events, never real content (diegosouzapw#12741) Validado numa worktree combinada com a onda de streaming desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 88/88 nos testes focados. Um stream que só emite eventos de ciclo de vida e heartbeat, sem conteúdo nenhum, é falha disfarçada de sucesso: o cliente espera até o timeout dele. Falhar rápido devolve o controle.
…2941) Validado numa worktree combinada com a onda de streaming desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 88/88 nos testes focados. Estender a rotação que já existe para 429 ao 403 de bloqueio geográfico é a generalização certa, e manter a rejeição de fingerprint (Cloudflare 1010) fora dela é o que impede a rotação de queimar todas as contas contra uma recusa que não é de egresso. Nota: os checkboxes de validação do corpo ficaram em branco, mas o diff traz dois arquivos de teste — vale marcar da próxima para o revisor não precisar conferir.
…n in-memory pending store (diegosouzapw#12854) Diagnóstico por captura de pacote em tráfego real, com o `400 previous_response_not_found` reassemblado do tcpdump três vezes no mesmo loop de tool-calling — isso é evidência, não hipótese. A causa é limpa: `detail_state` só vira `'ready'` depois de uma escrita fire-and-forget enfileirada num worker único, e o cliente já tem o id de resposta antes disso. Semear a ponte **antes do primeiro await** é o que faz a correção não custar latência. Revalidei sobre o tip: **19/19**, typecheck:core limpo, check-file-size OK. **Estava draft e eu marquei como ready.** Não havia gate declarado — nem RFC pendente, nem decisão de produto em aberto — e passou na validação; a diretiva permanente do dono para esta campanha é avaliar draft como qualquer PR e promover quando passa. Se a intenção era segurar por outro motivo, me avise que eu reverto. **Um conserto meu na sua branch.** O `typecheck:core` falhava com `TS2345` em `callLogs.ts:489` — e falhava **na sua branch sozinha**, não por interação com a onda; confirmei isolando. O call site fazia cast para `{ clientRawRequest?: unknown; clientResponse?: unknown }`, mais frouxo que o `ContinuationPipeline` que o parâmetro exige, e `unknown` não assina para os membros tipados. Exportei o `ContinuationPipeline` do próprio store e usei ele no cast, em vez de alargar o tipo do parâmetro: o contrato passa a ter um nome só, no lugar onde ele já vivia. **Sobre a sua Reviewer Note do Map sem limite de contagem:** concordo que vale registrar. Entradas pequenas com TTL de 60s auto-expirando não justificam sizing agora, mas se aparecer burst sustentado o sintoma será memória, não erro — e aí a nota está aqui. Também carreguei o rebaseline de `chatCore.ts` (6021→6026) e `stream.ts` (3080→3098), que a onda de streaming inteira faz crescer.
…nal (diegosouzapw#12639) (diegosouzapw#12988) * feat(api): hydrate memoryHits from the persisted history event `GET /api/a2a/tasks/[id]` falls back to the persisted history row once a task leaves the in-memory TTL window, and `reconstituteHistoricalTask` hard-coded `metadata: {}` — so the drawer's "Memory used" section vanished for any historical task, even though `executeA2ATaskWithState` had already written a `memory_hits` event with the hits. The fallback now reads that event: `data_json` is parsed and, when it yields at least one well-formed hit, exposed as `metadata.memoryHits`. The event itself is filtered out of `events` — it is observability, not a state transition, and without the filter it leaked into the timeline as a duplicate of the row's current state. Reading is defensive throughout, mirroring `DrawerMemory`'s own validation: the payload is caller-influenced and unvalidated end to end, so `JSON.parse` runs inside `safeJsonParse`, non-arrays are rejected, and each entry must carry `id`, `key`, `type` and `snippet` as strings (a non-string field would be rendered as a React child and take the drawer down). Malformed input degrades to `metadata: {}` and a 200 — never a 500. Refs diegosouzapw#12639 * fix(a2a): bound the memory recall with its own deadline collectMemoryHits() runs BEFORE the skill handler and had no deadline at all, so a slow memory backend delayed the start of every A2A task — the HTTP genericBackend alone defaults to a 30s timeout. The search now races a MEMORY_RECALL_TIMEOUT_MS (1500ms) deadline. Overshooting degrades exactly like any other recall failure: empty hits, a warn log, and the task proceeds normally (best-effort contract unchanged, nothing propagates). The deadline timer is cleared in a finally on BOTH paths so no handle is left holding the event loop open, and MemoryHitsDeps.timeoutMs makes it injectable so the tests cost milliseconds instead of 1.5s of wall clock. Refs diegosouzapw#12639 * fix(dashboard): carry conductor requirements and focus the repeated task The drawer's "Repeat" for a Conductor task dropped the runner/model pinning and left the operator staring at the finished run: - `hubTaskSchema` now parses the hub's `requirements` (`.catch(null)` so an odd shape never fails the whole task parse), and `ConductorTaskDetail` exposes `cli`/`model` (`null` when the hub sends none). - `repeatReqForConductor` carries `cli`/`model` when present and OMITS them otherwise — the route's Zod takes both as optional strings, so a `null` would 400. The two fields are independent. - `performAction` reads the response body once and returns it, so the repeat can report the CANVAS id of the created task (`task_id` / `data.id` / `result.task.id`, each with its node prefix). `OrchestrationPageClient` then refetches and focuses it via `?node=`; History keeps its current behavior. - `conductor-routes-auth.test.ts` covers the creation route through its `ROUTES` array; the duplicated source assertion left `conductor-create-route.test.ts`. Refs diegosouzapw#12639 * chore(a2a): follow-ups changelog Changelog fragment for the five items PR-C delivers from diegosouzapw#12639. The sixth item on the issue — an authenticated panel path for A2A task creation — stays deliberately out of scope and is recorded as such in a comment on the issue rather than silently dropped: the JSON-RPC endpoint accepts API keys only, and widening that endpoint's auth surface to serve a UI convenience is the operator's call, not the implementation's. Closes diegosouzapw#12639
… across flow surfaces (diegosouzapw#12378) (diegosouzapw#13203) * refactor(ui): move shared flow colors to the orchestration status tokens FLOW_EDGE_COLORS and TokenHealthBadge were pinned to the fixed dark-mode hexes in STATUS_HEX, so both rendered dark-theme green/amber/red on a light background. They now read the theme-aware --orch-status-{success,warning, error,muted} custom properties introduced in Fase 2. The dark values of those tokens are exactly the old hexes, so dark mode is unchanged and only light mode gains contrast. `idle` was already a CSS var, which is the precedent proving a var() resolves in a ReactFlow edge stroke. Five call-sites built translucent variants by concatenating an 8-bit alpha suffix onto the palette hex (`${FLOW_EDGE_COLORS.error}40`), which cannot work with a var(). They move to a new documented helper, flowColorAlpha(), that wraps color-mix() — the same approach orchStateBadgeBg() already uses in the orchestration model. Percentages mirror the old suffixes (20 -> 13%, 30 -> 19%, 40 -> 25%). STATUS_HEX stays exported as the dark-mode mirror; it now has no production consumer. globals.css needed no change — all five tokens already existed in both themes. The colour assertions in the topology, combo-live and design-grid suites were aligned to the tokens, never weakened: every hex equality became an equality against the corresponding var(). design-grid additionally now asserts each token is defined in BOTH themes. Refs diegosouzapw#12378 * refactor(ui): finish the status-token migration across flow surfaces Sweeps the five state hexes across the remaining flow surfaces, following D1: - ComboLiveStudio: active/error provider pills and the run-outcome tri-state. - CompressionCockpit / WaterfallInspector / IoNode: the savings readouts and the savings quality ramp (>=30 success, >=15 warning, else muted). - EngineNode: the same ramp, plus the running state, whose glow moved to flowColorAlpha — the literal #f59e0b40 suffix is invalid once the value is a var(). - WaterfallInspector: a skipped step now reads as muted rather than a bare grey hex. Deliberately NOT migrated, because they are categorical or brand palettes rather than state: STRATEGY_COLORS (routing-strategy hues), LAYER_COLORS (compression layer pills), the provider brand color in ProviderTopology, and IoNode's indigo/green input-output identity pair. The new test asserts both halves — what became a token AND what stays hex — so a later sweep cannot silently swallow a categorical palette. Refs diegosouzapw#12378
…egosouzapw#12876) Validado numa worktree combinada com a onda de dashboard/monitoring desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 130/131 nos testes focados — a falha restante é asserção de tempo de parede sob carga, verde 6/6 isolada. Um health que diz "falhou" sem dizer **qual** conexão obriga o operador a cruzar logs para achar o óbvio. Expor os ids das que falharam é o que transforma o endpoint em ferramenta de diagnóstico.
…iegosouzapw#12882) Validado numa worktree combinada com a onda de dashboard/monitoring desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 130/131 nos testes focados — a falha restante é asserção de tempo de parede sob carga, verde 6/6 isolada. Paginar e completar o stream em `GET /v1/models` é a correção certa para catálogo grande: um payload único que cresce com o número de providers vira timeout silencioso no cliente, não erro.
…n on one model 402 (diegosouzapw#12875) Validado numa worktree combinada com a onda de dashboard/monitoring desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 130/131 nos testes focados — a falha restante é asserção de tempo de parede sob carga, verde 6/6 isolada. Envenenar a conexão inteira por um 402 de **um** modelo é o erro clássico de granularidade em provider openai-compatible com múltiplos upstreams — derruba modelos que estavam saudáveis. Restringir ao modelo afetado é o comportamento correto, e é a mesma distinção que o guia de resiliência faz entre cooldown de conexão e lockout de modelo.
…instead of "(empty)" (diegosouzapw#12727) Validado numa worktree combinada com a onda de dashboard/monitoring desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 130/131 nos testes focados — a falha restante é asserção de tempo de parede sob carga, verde 6/6 isolada. "(empty)" para um nó de ferramenta ainda não resolvido é informação errada, não ausência de informação — o usuário lê como "não retornou nada". Spinner de pendente diz a verdade.
…egosouzapw#12937) Validado numa worktree combinada com a onda de dashboard/monitoring desta leva sobre `release/v3.8.51`: typecheck:core limpo, check-api-typecheck OK (289), check-file-size OK após rebaseline, 130/131 nos testes focados — a falha restante é asserção de tempo de parede sob carga, verde 6/6 isolada. Sentinela que não se explica (`—`, `?`) faz o leitor inventar a razão. Explicar no hover é metade; o `check:radar-sentinels` é a outra — sem o gate, a explicação apodrece na primeira coluna nova.
…iegosouzapw#12884) Um `ALL_TARGETS_SKIPPED` 503 que não diz qual janela esgotou é opaco justamente no momento em que o operador mais precisa saber. Alinhar os rótulos de janela AUTH com os da API de uso fecha a outra metade: dois nomes para a mesma coisa fazem o dashboard e o erro parecerem discordar. Revalidei sobre o tip: **6/6**, typecheck:core limpo, check-file-size OK. **Dois consertos meus na sua branch.** 1. `typecheck:core` falhava com `TS2345` em `comboAttemptLoop.ts` (linhas 130 e 416): o `QuotaSkipTarget` declarava `connectionId?: string`, mas o `ResolvedComboTarget` carrega `string | null` para alvo não-pinado. Alarguei para `string | null` no tipo de diagnóstico em vez de estreitar o call site — o módulo só **lê** o campo e a linha 29 já narrowa com `typeof === "string"`, então null não custa nada ali. Isso apareceu porque o `comboAttemptLoop` mudou de forma no diegosouzapw#12746/diegosouzapw#12811, mergeados nesta mesma campanha depois que você cortou a branch. 2. O `roundRobinCombo.ts` foi de 1198 para 1205 e cruzou o teto de 1200 para arquivo novo. Congelei com justificativa: o arquivo já nasceu em 1198 quando o diegosouzapw#12811 o levantou de dentro do `combo.ts`, e os diagnósticos em si vivem no `quotaSkipDiagnostics.ts`, sob o cap. Registrei que a próxima extração natural é o corpo do attempt loop, mas que ele acabou de ser movido e deve assentar antes de ser cortado de novo.
…iegosouzapw#12715) Fila sem teto que segura a request seis minutos até o cliente abortar é pior que 503 imediato: consome slot, mascara a saturação e ainda entrega erro no fim. Um orçamento `maxWaitMs` por conexão compartilhado entre gate, slot padrão do provider e fila do Bottleneck é a forma certa — o teto tem que ser um só, senão cada camada espera o seu. O `max(perConn, upstream)` no `executionMaxWaitMs` é o detalhe que evita a correção matar request em voo, que seria trocar um defeito por outro. Registro a atribuição: você manteve o diegosouzapw#12635 aberto para o @Tushar49 e creditou a percepção dele (providers lentos precisam de 2min→10min por conexão) enquanto adiciona o encanamento que faltava. É o jeito certo de construir sobre PR de outra pessoa sem tomar o crédito. Sobre o `npm run lint` desmarcado com a nota do eslint quebrado no ambiente: deixar em branco e explicar vale mais que marcar sem ter rodado. Rodei aqui: limpo. Revalidei sobre o tip: **13/13**, typecheck:core limpo, check-file-size OK. O `file-size-baseline.json` conflitou com os rebaselines desta campanha — resolvido aditivamente, JSON revalidado com `json.load`.
diegosouzapw
merged commit Sep 10, 2026
955b28e
into
diegosouzapw:release/v3.8.51
4 of 7 checks passed
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…diegosouzapw#12717) Uma conversa que nunca chegou a parada limpa e não sinaliza nada é o pior estado possível de UI: indistinguível de uma que terminou. O incidente que você cita no comentário do teste — stream pesado em reasoning estourando o cap do coletor no meio, deixando a conversa presa sem sinal — é exatamente o caso que justifica o badge. Separar `resolveTurnCompletionState` de `resolveConversationStalledState` também está certo: `tool_call_pending` é um estado legítimo em voo, não uma conversa travada. Revalidei sobre o tip: **29/29**, typecheck:core limpo. **Nota de integração.** O `tests/unit/responses-continuation-store.test.ts` conflitou com o diegosouzapw#12854, que anexa a própria bateria ao mesmo arquivo. Reconstruí o arquivo como append limpo — versão do tip mais o seu bloco de 184 linhas, verificado por `esbuild` antes de rodar. Registro por que importa: na primeira tentativa eu apenas retirei os marcadores de conflito, e isso enfiou os seus testes **dentro** de um objeto literal não terminado do diegosouzapw#12854. Compilava como erro de transform, não como conflito — só apareceu ao rodar. Resolver JSON e teste "aditivamente" sem verificar a sintaxe depois é armadilha; ficou a lição.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What Problem This Solves
A conversation whose stream genuinely failed mid-flight (blew past the SSE
collector's own truncation cap, or a tool call the agent never followed up
on) shows no visible signal on
/dashboard/conversations— it just sitsthere looking identical to any healthy conversation.
Why This Change Was Made
Live incident: a reasoning-heavy stream blew past
createStructuredSSECollector's own event-count cap mid-stream (3013 SSEevents, 1487 dropped). The stored
clientResponsenever reached"completed"(_truncated: true,summary.statusstuck at"in_progress",output: []) — the client never got a valid reply, andnothing surfaced this.
A bare unanswered tool call is deliberately not treated the same way —
it's the completely normal shape of a mid-conversation turn seconds after it
lands, before the agent runtime sends the next request with the tool's
result. Flagging that immediately would false-positive on every healthy
in-flight tool-calling conversation, so it gets a 5-minute grace period
before counting as a stall.
Fix
resolveTurnCompletionState(responsesContinuationStore.ts): a pure,permanently-cached function of one call-log's own already-persisted
artifact, returning
"stop" | "tool_call_pending" | "incomplete" | "unknown". Same immutability argument and cache shape as the existingsibling
isGenuineContinuationTurn— an artifact never changes oncedetailStateflips to"ready".resolveConversationStalledState: combines that with the liveisActive/lastSeenAtfacts (never cached) to decide the badge. Nevertrue while
isActive.annotateConversationRowshared by both thelist and single-conversation routes — no new API surface, no schema
change.
/dashboard/conversations, next to the existing"continuation" badge.
User Impact
A conversation that silently failed (truncated stream, or a tool call
nobody ever answered) is now visibly flagged instead of looking identical to
a healthy one. No behavior change for any conversation that ended cleanly.
Evidence
tests/unit/responses-continuation-store.test.ts: 10 new cases covering allthree completion states (
stop/tool_call_pending/incomplete,including the exact truncated-collector shape from the live incident), the
5-minute grace-period boundary in both directions,
isActivealwaysoverriding to
false, and a cleanstopnever flagging regardless ofelapsed time. Full file: 23/23 passing.
Verified live against the real flagged conversation from the incident:
/api/conversationsnow returnsisStalled: truefor it (confirmed via adirect authenticated request against the running dev deployment).
eslint/tsc --pretty false -p tsconfig.typecheck-core.jsonclean on alltouched files.