fix(stream-recovery): log recovery traces with continuation attempt - #13650
Merged
diegosouzapw merged 84 commits intoSep 15, 2026
Merged
diegosouzapw merged 84 commits into
diegosouzapw merged 84 commits into
Conversation
maxmad64bis
force-pushed
the
fix/recovery-trace-wiring
branch
5 times, most recently
from
September 14, 2026 23:00
f9180f6 to
7a7f9b3
Compare
maxmad64bis
force-pushed
the
fix/recovery-trace-wiring
branch
from
September 14, 2026 23:07
7a7f9b3 to
5c73397
Compare
…apw#13286) (diegosouzapw#13418) Chains `npm run test:unit:serial` onto the plain `npm test` script, so the quarantined serial suite is no longer skipped when contributors run `npm test` (diegosouzapw#13286). The existing `test-serial-quarantine` guard now covers `test` too. Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…uzapw#13273) (diegosouzapw#13401) Masks every `*-api-key` header spelling (`x-goog-api-key`, `api-key`, `xi-api-key`) in the request-log pipeline by matching on the hyphen-compacted key, and adds `xi-api-key` to the payload redaction set (diegosouzapw#13273). The new test drives the real `createRequestLogger` / `protectPayloadForLog` paths. Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…iegosouzapw#13403) Fixes the stale `jina-ai/` prefix in the second Jina assertion of `models-catalog-route.test.ts`. The catalog emits `jina/…` (the first Jina test already expected it), so this clears a base red on the release branch. Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…iegosouzapw#13308) (diegosouzapw#13404) Prunes `db_backups/` after the health-check-repair `VACUUM INTO` snapshot (diegosouzapw#13308). That path never ran retention, so every restart of a healthy database added another full-size copy. It now uses the same `DB_BACKUP_MAX_FILES` / `DB_BACKUP_RETENTION_DAYS` limits as `backup.ts`, importing `backupRetention` directly to avoid the import cycle. Maintainer additions: merged the current release branch and rebaselined `src/lib/db/core.ts` 1745→1767 in `file-size-baseline.json` with a dated annotation (legitimate growth). Chosen over diegosouzapw#13540, which did the same with a fire-and-forget dynamic import. Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…ction (diegosouzapw#12776) (diegosouzapw#13406) Carries `context_length` through the MCP `list_models_catalog` projection when the upstream model record has it, and omits it otherwise (diegosouzapw#12776). Covered by the new case in `mcp-model-catalog.test.ts`. Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…#13409) "Test all models" no longer sends image/music/video generation-only models through a chat completion, which triggered real billable generations (diegosouzapw#13376). `detectTestKind` flags them via `isNonChatGeneration` and `runSingleModelTest` skips them; chat+image models, embeddings and models without metadata are unaffected (8 new cases plus updated `model-test-runner` shapes). Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…iegosouzapw#13412) Adds CJK quota-exhaustion messages (GLM/z.ai 5-hour window, Kimi, Qwen/DashScope, MiniMax, plus Japanese and Korean phrasings) to the 429 quota classifier. They are now `quota_exhausted` (cooldown + failover) instead of a transient `rate_limit` retry loop (diegosouzapw#13194). The patterns go into the recoverable `QUOTA_PATTERNS` only, not the terminal set, and a guard case keeps the Chinese transient "请求过于频繁" as `rate_limit`. Chosen over diegosouzapw#13536. Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…otAll regex (diegosouzapw#13413) Eval runner: a case whose upstream call failed can no longer score as passed when the grading regex happens to match the error text (diegosouzapw#13137), and regex grading compiles with dotAll so `.` spans newlines in multi-line answers, for both string and `RegExp` patterns (diegosouzapw#13138). Chosen over diegosouzapw#13542 / diegosouzapw#13530, which each covered half. Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…on (diegosouzapw#13414) The eval runner sends `x-omniroute-compression: off` and `x-omniroute-no-memory: true`, so cases measure the model rather than injected output styles or retrieved memory (diegosouzapw#13139). Both headers are the existing per-request opt-outs honored by `chatCore` (`open-sse/handlers/chatCore/headers.ts`). Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…iegosouzapw#13411) The media-providers web-search example card now sends `provider` in the request body. `POST /v1/search` selects the provider from `body.provider`, so the card was silently hitting the default provider (diegosouzapw#13245). Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…osouzapw#13410) De-duplicates the `agy` and `antigravity` model catalogs into `antigravitySharedModels.ts`, with an explicit per-surface add/remove delta (diegosouzapw#12724). Verified behavior-neutral: every exported catalog constant of both modules serializes byte-identically before and after (201-line JSON dump). Maintainer fix: `buildSurfaceCatalog` returned `readonly { id: string; [k: string]: unknown }[]`, which erased the element type. `name` became `unknown`, and the `RegistryEntry.models` assignments failed with 4× TS2322 under `check:open-sse-typecheck` (not covered by `typecheck:core`). It is now generic (`<T extends { id: string }>`), so the literal model shape flows through. Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…ifting at messages[0] (diegosouzapw#13425) (diegosouzapw#13427) Memory injection merges into an Anthropic-shaped top-level `system` field (string or text-block array) instead of prepending a `role: "system"` message at `messages[0]`, which Anthropic rejects with a 400 (diegosouzapw#13425). Handled on both the system-first path (xiaomi-mimo) and the general path. Chosen over diegosouzapw#13549, which covered only the string case. Maintainer fix: widened `ChatRequest.system` to `string | Array<{ type; text?; … }>`. Assigning the block array to the `string`-typed field raised 2× TS2322 under `check:open-sse-typecheck`. Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…3369) (diegosouzapw#13433) The CLI readiness budget is configurable through `OMNIROUTE_READY_TIMEOUT_MS` or `omniroute serve --ready-timeout <ms>` (default unchanged at 60s). The timeout warning prints the budget it actually used and suggests a larger value, for slow cold starts such as Windows (diegosouzapw#13369). Documented in `ENVIRONMENT.md` and `TROUBLESHOOTING.md`, with 11 resolver cases. Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…iegosouzapw#12779) (diegosouzapw#13340) Adds the Microsoft 365 Copilot (BizChat) provider guide for the existing `copilot-m365-web` provider (alias `m365copilot`, `open-sse/config/providers/registry/copilot-m365-web`), including the WebSocket credential capture flow, and links it from the docs index and the web-cookie guide (diegosouzapw#12779). Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…bility (fixes diegosouzapw#13467) (diegosouzapw#13539) Adds a `DISABLE_IOREG_STRATEGY=1` escape hatch to the macOS `ioreg` strategy in `getMachineIdRaw()`, so `machineId.test.ts` can reach the fallback strategies on darwin instead of always resolving the real hardware UUID (diegosouzapw#13467). Maintainer note: this branch also carried the "force passed=false when upstream call failed" eval commit, which already landed in diegosouzapw#13413. The branch was reset to the current release tip plus only the ioreg commit (authorship preserved) so the squash carries just this change. Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…failure (diegosouzapw#13546) Combo attempt call logs are keyed on the per-attempt `traceId` instead of the shared `pendingRequestId` (diegosouzapw#13481). Every attempt of a failed-over combo reused the same id, so the successful member hit the `call_logs` UNIQUE constraint and was silently dropped: the dashboard showed only the failed steps. Maintainer additions: carried your regression test from diegosouzapw#13526 (`combo attempt uses traceId as the log id, not pendingRequestId`, two attempts sharing one request id both persist) into this cleaner branch. Removed the now-unused `pendingRequestId` destructure (`no-unused-vars`) and added a short comment at the call site. Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…ection (diegosouzapw#13547) No-auth providers now honor a recorded model-only lockout before `getProviderCredentials` hands back the synthetic `noauth` connection (diegosouzapw#13483). That early return skipped the per-connection status pass, so a `model_capacity` lockout was recorded but never enforced, and every request re-tried the locked model for a wasted upstream round-trip before failing over. Maintainer additions: carried your `tests/unit/noauth-model-lockout.test.ts` from diegosouzapw#13527. 3 of its 4 cases fail on the release tip without the fix and pass with it. Rebaselined `src/sse/services/auth.ts` 3542→3556 in `file-size-baseline.json` with a dated annotation. Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…gosouzapw#13523) Lossy compression engines leave `<system-reminder>`, `<instructions>` and `<project-instructions>` envelopes byte-identical. Agentic CLIs inject these into user messages, and compressing them inverted negations, dropped emphasis and broke the tags (diegosouzapw#13453). Maintainer fixes: (1) the new test imported `../../open-sse/…` from `tests/unit/compression/`, which does not resolve, so it had never run. With the path fixed, the main case failed: the envelopes were appended to the built-in list, so fenced code and inline code inside the reminder had already become sentinels, and `replacePattern` skips any match containing one. The envelopes now form a region pass right after frontmatter, before fenced-code extraction; all 4 cases and the full compression suite (1,518) pass, and the compression budget gate reports no regression. (2) Dropped an unused `tombstoned` binding. (3) The branch also carried an unrelated `feat(sveltekit)` commit (`apps/web/**`), so it was reset to the release tip plus only this commit, authorship preserved. Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…stic (diegosouzapw#13524) The ultra heuristic no longer prunes polarity and modality words (`never`, `always`, `not`, `must`, `should`, `do`/`does`/`did`, contractions…). They now score 1.0 instead of the 0.1 stopword score, so "must never be deleted" can no longer collapse into "must deleted". It also collapses only runs of spaces and tabs, keeping the newlines that carry bullets, headings and fences (diegosouzapw#13454). Maintainer fixes: (1) the new test imported `../../open-sse/…` from `tests/unit/compression/`, which does not resolve, so it had never run. It now passes with the rest of the compression suite (1,526 cases). (2) Dropped the unused `STOPWORDS` import. (3) `check:compression-budget` flagged the expected trade-off: ultra tokens-per-task goes prose 92→97, tool-output 116→117, json 117→126, the cost of keeping meaning-bearing words and line structure. The baseline was refreshed with `--update`, which also records a pre-existing caveman tightening (129/127/160 → 119/114/136). (4) The branch also carried the unrelated `feat(sveltekit)` commit and the diegosouzapw#13523 commit, so it was reset to the release tip plus only this commit, authorship preserved. Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…iegosouzapw#13457) (diegosouzapw#13525) `cavemanConfig.preservePatterns` regions are captured before the built-in patterns. Previously the built-ins ran first and turned inline code, URLs and CONST_CASE identifiers inside the user region into sentinels, and `replacePattern` silently skipped the user match (diegosouzapw#13457). Maintainer adaptation: after diegosouzapw#13523 moved the instruction envelopes into a region pass that runs before fenced-code extraction, the user patterns now open that same pass. User regions win over the built-in envelopes, and fenced code inside a user region stays with the region too. This keeps the "run user patterns first" semantics of your change. Also fixed the new test's `../../open-sse` import path (it had never run) and three unused `tombstoned` bindings. Reset to the release tip plus only this commit (the branch carried ~2.6k unrelated commits), authorship preserved. All 5 new cases pass; full compression suite 1,531/1,531; budget gate clean. Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…30s (diegosouzapw#13553) Raises the default request-queue wait (`RATE_LIMIT_MAX_WAIT_MS`) from 15s to 30s, so bursts behind a rate-limited provider queue a bit longer before being rejected (diegosouzapw#13504). The env var still overrides it, and `executionMaxWaitMs` (the 10-minute execution backstop) is unchanged. Approved by the maintainer as a policy change. Maintainer additions: updated the documented default in `.env.example` and `docs/reference/ENVIRONMENT.md` to match. Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green. Thanks @KooshaPari!
…dget 429s (diegosouzapw#13606) `token_limit_exceeded` joins `REQUEST_SCOPED_UPSTREAM_ERROR_CODES`: chatCore's local Tier-2 check answers 429 with that code, but `shouldSkipConnDisable` did not know it and cooled a healthy connection down for a request-sized problem. Combo exhaustion treats it as request-scoped too. Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
…ceeds its time bound (diegosouzapw#13438) A cold `/v1/models` catalog build that exceeds its time bound now answers 503 with `Retry-After` instead of a 500, and the timed-out build stays joinable so the next retry does not start another cold build. Seven cases; the first fails on the tip (500 → 503). Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
…gosouzapw#13279) Combo test probes are aborted when the dashboard client disconnects (`AbortSignal.any` over the route's own timeout and `request.signal`), instead of running to completion for nobody. Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
diegosouzapw#13281) `error_type` becomes a versioned vocabulary: a failure the classifier cannot place is stored as `unknown` instead of NULL, and any stored value outside the vocabulary reads back as `unclassified` in the breakdown. Maintainer rework before merge (kept the idea, no default behavior change): - `PROVIDER_ERROR_TYPES` is `as const`, so `ErrorTypeContract` is a real union and the classifier functions return the narrowed type. - The constants moved above the JSDoc that documents `getErrorTypeBreakdown`; the vocabulary SQL is built on first use so an import cycle cannot read it before it exists. - Legacy NULL rows keep landing in the `pre_migration`/`unclassified` bucket without vanishing or double counting, and the log export / BigQuery row pass both NULL and `"unknown"` through unchanged — both covered by tests. Duplicate assertions removed; every seeded row is cleaned up in `finally`. Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
) A write-boundary guard for `error_type`: `toStoredErrorType()` validates what `saveCallLog` stores against the vocabulary (Zod enum built once), as defense in depth on top of diegosouzapw#13281. Maintainer rework before merge (kept the idea, no default behavior change): - Dropped the redundant `SCHEMA_SQL` column (migration 158 already creates it) and the string-absence "migration 177" test; the real `PRAGMA table_info` test is back. - Restored diegosouzapw#13281's changelog fragment, which this branch had deleted, and renamed this PR's own fragment to `13441-error-type-write-guard.md`. - The guard is now exercised for real: the exported function is tested with out-of-vocabulary values and an end-to-end drift test that changes a classifier family at runtime. Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
…ted as secrets (diegosouzapw#13729) The labels wvxc-route-401/wvxc-route-500 sat right after a key*.id argument and cleared the gitleaks generic-api-key length and entropy floors; renamed to route401/route500 with a docblock stating the measured rule. Test-only; .gitleaks.toml untouched. Reviewed by 3 rounds of /omni-code-review (37 agents).
…w#13641) Search stats and recent searches stop surfacing ghost rows: NULL and `-` providers are always hidden, and, behind the new `SEARCH_STATS_HIDE_DELETED_CONNECTIONS` flag (default off), traffic of a keyed provider whose connection was deleted is hidden too. Totals use the same guard as the per-provider rows, so they always agree. Maintainer rework before merge (kept the idea, no default behavior change): - Keyless providers from the search registry (`duckduckgo-free`, `searxng-search`, anonymous `context7`) and providers served through a credential fallback (`perplexity-search` on a `perplexity` key) stay visible in both modes — the original filter dropped them because they have no `provider_connections` row. - Tests use real registry ids and cover flag off (historical stats) and flag on, including the analytics route; diegosouzapw#13281's changelog fragment restored. Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
… and make the guard find them (diegosouzapw#13436) Fixes the production build break from `node:fs` reaching client bundles (`oauth.ts → cursorAgentCliVersion.ts` through the codebuddy-cn registry) and widens the client-bundle guard so it finds any Node builtin, not just the one that broke. Maintainer rework before merge (kept the idea, no default behavior change): - The guard was 11× slower (3.8s → ~40s) because resolved edges were not cached; with resolved edges and per-file verdicts cached it runs in ~4.6s. - Bare builtins that Next's client build polyfills (`path`, `os`, `crypto`, `buffer`, … from Next's own `resolve.fallback` list) are allowed consistently; `node:` imports are always flagged; a drift test fails if Next stops polyfilling an allowlisted name. Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
…uzapw#13605) Proxy credentials containing a literal `%` no longer throw `URIError`: every `decodeURIComponent` on proxy user/password is guarded. Maintainer rework before merge (kept the idea, no default behavior change): - HTTP proxies still failed because undici's `ProxyAgent` decodes the credentials itself; the dispatcher now builds undici's `Basic` token with the safe decoder and passes it as `token`, so a literal `%` works there too. - The three remaining unguarded sites (`mappers.ts`, `proxySubscription/parse.ts`, `subscriptionService.ts`) are guarded; tests run the real `createProxyDispatcher` against a local HTTP CONNECT proxy and a local SOCKS5 server that record what they received. Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
…osouzapw#13280) Seven routing/quota caches (quality states, account buckets, quota-fetcher/saturation/header caches, learned rate limits) sit behind a shared bounded map with LRU/TTL eviction instead of growing without bound. The learned-limits cap of 200 that the tip declared was never enforced. Maintainer rework before merge (kept the idea, no default behavior change): - Eviction logging goes through the project logger, aggregated (first eviction, then one summary line per minute per map) instead of a `console.warn` per eviction. - `refetch-lazy` and `hard-expire` behaved identically and are collapsed into `ttl`; protected entries (saturated account buckets, evaluator quality scores) are never evicted; caps raised to 2048–4096 so normal deployments never evict, with tests showing 300 learned limits and 600 cached entries all kept. Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
The WAL busy counter survives restarts: it is persisted in `key_value` and restored at boot, so the health output no longer resets to zero after every restart. Maintainer rework before merge (kept the idea, no default behavior change): - `recordBusy()` no longer writes synchronously on the contended path (with `busy_timeout = 2000` that could block the event loop for up to 2s); it accumulates in memory and `flushBusyTotal()` upserts the delta on a clean passive/TRUNCATE tick or best-effort at stop. - The boot wiring is tested for real: a child Node process drives `startWalMaintenance` against a real SQLite file (restore at boot, zero writes while busy, one flush at stop, restore after restart). Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
Dependabot diegosouzapw#214 (GHSA-vwc7-r8mq-g2x9 / CVE-2026-76845, moderate): adm-zip 0.5.9–0.6.0 follows a symlink that already exists inside the extraction root and writes through it, outside the root. The advisory still reports `first_patched_version: null`, but 0.6.1 (published after the advisory) is the fix — `util/utils.js` gains `assertPathSafe`, which walks every path component below the root with `lstat` and throws on a symlink; `extractAllTo` calls it before every write. Verified by diffing the two tarballs. Reach in this repo: adm-zip is pulled only by `onnxruntime-node` (an optionalDependency, itself pinned by override) and used only in its install script to unpack the vendor's own runtime binary. No request path touches it. The override already existed at ^0.6.0 (PR diegosouzapw#7732, the previous adm-zip CVE); this just raises the floor. Lockfile moves 0.6.0 → 0.6.1, nothing else.
…credential catalog (diegosouzapw#13744) GHSA-r4q7-7f24-m29p. `CREDENTIAL_PATTERNS` (open-sse/utils/credentialPatterns.ts) is the single catalog iterated in order by both the opt-in credential-masker guardrail and the public error sanitizer. It had no entry for Groq (`gsk_`) or xAI (`xai-`), and only knew the exact 48-char OpenAI `sk-` form. Measured on the release tip before this change: | Shape | public sanitizer | guardrail | |------------------------------|------------------|-----------| | Groq gsk_ + 52 | LEAK | LEAK | | xAI xai- + 80 | LEAK | LEAK | | DeepSeek sk- + 32 hex | redacted | LEAK | | sk- + 20/36/40/51 (not 48) | redacted | LEAK | The public path already caught every `sk-` shape through STRONG_CREDENTIAL_TOKEN, so the advisory's "both layers" framing only holds for gsk_/xai-; for the sk- family the exposure was the guardrail. Adds `groq` and `xai` after `anthropic_alt`, and a generic `openai_compatible` `sk-` fallback as the LAST entry. Ordering matters: both consumers replace as they iterate, so `openai_proj`, `openai` and `anthropic*` stamp their specific label first and the fallback only sees shapes nothing else claimed. The lookbehind mirrors STRONG_CREDENTIAL_TOKEN so `risk-…`-style words do not match. All three regexes are a fixed prefix plus one bounded character class — linear, no nested quantifiers. Tests are red-first: the new guardrail cases (bare / sentence / JSON-body contexts per shape, plus label-ordering and negative cases) and the catalog coverage array in error-sensitive-redaction both failed on the tip. Follow-ups deliberately left out of scope: `tskey-auth-` (Tailscale) was never in the catalog, and the guardrail does not decode `\uXXXX` escapes the way the public path does.
diegosouzapw#13748) GHSA-34rg-3pqj-35g9. `fetchRemoteImage()` defaults to `getProviderOutboundGuard()` — the OPERATOR outbound policy, local-first by design so self-hosted providers on loopback/LAN keep working. Since diegosouzapw#11062 added the `block-metadata` middle tier, a default install resolves to that mode: the string check only rejects 169.254/16 and the IMDS hostnames, and the DNS validation step is skipped entirely (it only runs under `public-only`). Three sinks feed that default with CALLER input, so a request body could make the server fetch `http://127.0.0.1:…` or any RFC-1918 host and forward the bytes upstream: - imageGeneration.ts `resolveImageSource()` — `image_url`, `mask_url`, message parts - imageUpscale/shared.ts `resolveUpscaleImageSource()` — 14 body aliases, `provider_options.*`, message parts (Stability, Topaz) - visionBridgeHelpers.ts `fetchRemoteImageAsDataUri()` — chat `image_url` parts inlined into the vision self-call plus the NanoBanana result-URL download, which is upstream-supplied rather than OmniRoute-controlled. Same trust confusion as GHSA-3f8g / GHSA-j7j4 on the search base URL: operator config and caller input must not share a guard. Each site now passes `guard: "public-only"` (string check + DNS validation of every answer), matching the siblings that already did it right — embeddings, the audio bridge and the AI Horde result download. `pinDns` is set only on the vision bridge. The other three sites use `globalThis.fetch`, and connection pinning would swap that for a raw undici fetch — the same reason the AI Horde site leaves it off. On the vision bridge a `fetchImpl` is injected, so `pinDns` there validates every DNS answer but cannot pin the connection; commented in place. Blind SSRF rather than full read: the bytes go upstream or into the vision self-call, not back to the caller — but the status oracle and upstream exfiltration are real. Tests are red-first — per sink, `http://127.0.0.1:1/x.png` and `http://192.168.1.50/x.png` are rejected with the injected fetch never called, and a public host whose DNS resolves to a public IP still downloads.
… records and anonymous listing (diegosouzapw#13749) GHSA-2jm2-mpx8-6523 and GHSA-m3hp-hq9g-fpmv, one root cause. `getApiKeyRequestScope()` never rejects: with the default REQUIRE_API_KEY=false the client-api policy admits both a missing and an invalid bearer as anonymous, and the scope comes back `{ apiKeyId: null, isSessionAuth: false }`. The `/v1/files` and `/v1/batches` routes then treated "null" as permissive in two different ways: - GHSA-m3hp — the list routes coerced `apiKeyId || undefined`, and the DB layer reads `undefined` as "no owner filter", so an anonymous or invalid-bearer caller got every tenant's file and batch metadata, the same unfiltered view as the operator's dashboard. - GHSA-2jm2 — the single-record checks were `record.apiKeyId !== null && …`, so a record with no owner short-circuited to "allowed" for any caller: read, download raw content, delete, cancel, or use as a batch input. Null-owner records are common — every dashboard-session upload, and every batch output file inheriting a session batch's owner, which carries model responses. `api_key_id` has existed since the table was created (migration 028), so a null owner is not a legacy row; it is an unattributable write. No doc described it as shared — API_REFERENCE says files are scoped per key — and batch_api.test.ts pinned the by-id exposure as expected behaviour. One rule now, in `_helpers/apiKeyScope.ts`: - `canAccessOwnedRecord(scope, owner)`: a dashboard session is the instance operator and may act on any record; an API key acts on its own records only; a null owner is denied to every non-session caller. Applied to files GET / DELETE / content, batches GET / DELETE / cancel, and the batch-create input-file check. - `resolveListScope(scope)`: an explicit union for list/count reads — scoped to the presented key (a key wins even alongside a session cookie), instance-wide only for a session without a key, and 401 otherwise, including for a bearer that does not resolve to a key. There is no default that widens a read. This follows the GHSA-wvxc shape already used by the delete-completed sweep. Behaviour change: the anonymous upload → batch → download flow no longer works without an API key, because a null owner cannot be attributed. Subsumes diegosouzapw#13683: it moved `scopeCheck` into the shared helper so a session can cancel any batch — kept, and its test ported — but it also kept null-owner records open on the premise they predate ownership tracking, which migration 028 contradicts. Tests are red-first. batch_api's by-id case is flipped to 404 with a negative assertion; batch-deletion-route-logic now imports the real helper instead of a local copy that had silently diverged from production; the two integration tests present a real key, since their subject is limits and rate logging, not auth. Co-authored-by: Markus Hartung <mail@hartmark.se>
…LOCAL_ONLY (diegosouzapw#13745) GHSA-35fw-cv32-2373 and GHSA-jx89-f37j-pq89 — the same defect class as /api/acp/agents (GHSA-hf57): a route whose handler chain spawns a host process was classified Tier 3 MANAGEMENT only, and requireManagementAuth() waives auth when requireLogin=false. Hard Rules diegosouzapw#15/diegosouzapw#17 require the LOCAL_ONLY gate, which runs on the stamped real peer before any auth check. cli-tools (GHSA-35fw): 14 routes reach getCliRuntimeStatus() -> locateCommand() -> runProcess("sh", ["-c", 'command -v -- "$1"']) -> spawn(), exactly like their six gated siblings (forge/grok-build/jcode/qwen/omp/letta-settings): all-statuses, status, and the claude/cline/codewhale/codex/crush/deepseek-tui/ droid/kilo/openclaw/pi/smelt-settings routes. The advisory counted 13; it missed /api/cli-tools/detect, which is heavier — detectAllTools() runs execFile(binary, ["--version"]) and execFile("which") per tool. skills (GHSA-jx89): POST /api/skills/install stores the request's handlerCode verbatim as the skill handler with no allowlist, so a value equal to a built-in name (execute_command / eval_code) aliases the real sandboxed built-in; POST /api/skills/executions then runs it. The sandbox is a real container, but the spawn is transitive, which is why the 6A.8 source scan never flagged it. Entries are exact paths, not a /api/cli-tools/ blanket prefix: apply, backups, config, guide-settings, hermes-agent-settings, keys, logs, openclaw/auto-order and codex-profiles do not spawn and remote dashboards use them. All 16 are mirrored into SPAWN_CAPABLE_PREFIXES (no manage-scope bypass) and added to the route-guard-membership roots so the gate enforces them from now on. Functional trade-off, same one already accepted for grok/forge/jcode/qwen: a dashboard served through a tunnel no longer shows the CLI Tools status badges. Tests are red-first. Two existing negative controls pointed at routes that turn out to spawn (/api/cli-tools/all-statuses, /api/skills/install); they now point at routes that genuinely do not (/api/cli-tools/config, /api/skills/marketplace, /api/skills/skillssh/install), so the non-over-gating assertions are kept.
…ouzapw#13645) The provider-page Free badge can require a free tier the provider actually honors: behind the new `FREE_BADGE_REQUIRES_PROVIDER_FREE_TIER` flag (default off) the display-name "free" heuristic, non-boolean `free` fields and `:free` suffixes on registered providers without a documented free tier no longer light the badge. With the flag off the historical rule is unchanged. Maintainer rework before merge (kept the idea, no default behavior change): - `:free` models on free-tier providers and on compatible nodes (OpenRouter-style endpoints) keep the badge in both modes — the original change dropped them. - Test D derives its provider set from `FREE_MODEL_BUDGETS` instead of a hard-coded allowlist; a vitest render covers both sections with the flag endpoint on, off and erroring. Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
…apw#13646) Both settings routes use the shared `isSocks5ProxyEnabled()` reader instead of a copied check (identical logic, no behavior change). Maintainer rework before merge (kept the idea, no default behavior change): - The source-grep tests were replaced by behavioral tests of both routes across the flag on/off matrix (`GET /api/settings/proxies` reports `socks5Enabled`; `PUT /api/settings/proxy` accepts socks5 or returns 400). Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
…owns (diegosouzapw#13440) Daily-quota lockouts on the non-TPD path honor the provider's configured daily-reset clock (`dailyQuotaResetTimezone`/hour) in combo routing instead of the host's midnight. Maintainer rework before merge (kept the idea, no default behavior change): - The process-lifetime clock cache is gone: the clock is resolved on each failure through the already-TTL'd `getCachedProviderNodes`, so a timezone change takes effect without a restart and a DB error is never cached as `{}` forever. - Round-robin combos are threaded too (the PR left them out); an option on `recordModelLockoutFailure` that could never run was removed; tests prove both call sites pass the configured clock. Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
…ons (diegosouzapw#13671) Fixes the DST-gap bug in `nextDailyResetAtMs`: a reset hour that does not exist on the transition day landed one hour early (New York 02:00 came out as 01:00; Havana/Santiago midnight as 23:00 the day before). The walk across the gap is bounded to one day and uses a cached formatter. Maintainer rework before merge (kept the idea, no default behavior change): - Dropped the 24h clamp in `getMsUntilTomorrow` (on a 25h fall-back day 24.5h is the correct wait; clamping expired the lock 30 minutes early) and the unreachable `ms <= 0` branch, with their tests; characterization tests pin ordinary and fall-back days. Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
…in path (diegosouzapw#13672) Behind the new `RETRY_AFTER_PROVENANCE_ENABLED` flag (default off): `unavailableResponse` omits the synthetic `Retry-After: 1` when there is no real retry signal, marks `retry_after_provenance` on its bodies, and both combo drain readers parse prose retry hints from plain-text bodies too. With the flag off, headers and bodies are exactly as before. Maintainer rework before merge (kept the idea, no default behavior change): - A past `Retry-After` date is no longer labelled as an upstream signal with `Retry-After: 1`; non-JSON bodies (HTML 502 pages) log at debug instead of warning on every request. - The provenance claim is narrowed to responses built by `unavailableResponse`, documented in the flag row. Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
…ouzapw#13439) Behind the new `PROTECTED_PRIORITY_INFRA_502_ENABLED` flag (default off), protected-priority combo stops caused by provably non-quota infrastructure (provider circuit open, predictive-TTFT latency) surface as 502 instead of a quota-looking 503. Maintainer rework before merge (kept the idea, no default behavior change): - The original branch made 502 the default for every stop, including model lockouts and cooldowns, and removed the diegosouzapw#8133/diegosouzapw#1731 provider-wide skip for 401/5xx without a connection id; both are restored with their regression tests untouched. - Nineteen cases cover eight gate causes plus predictive latency, flag off and on. Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
…rs plus fire-and-forget async (diegosouzapw#13614) Stale provider-pin (`clearStaleLKGP`) clears are no longer silent: the fire-and-forget promise carries a `.catch` that warns with combo, comboId and executionKey, and a `check:routing-error-guard` npm script keeps the inventory of swallowed catches in the routing hot path from growing. Maintainer rework before merge (kept the idea, no default behavior change): - The awaited DB writes in the fallback loop were reverted (they added latency and SQLite lock exposure on every skip); the clear stays non-blocking. - The guard keys its allowlist by file + normalized catch body instead of line numbers (the PR's version broke on any edit) and is wired as an npm script only, not in CI; the unused stats counters were dropped. Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
…d off-by-default flag (diegosouzapw#13633) Behind `STREAM_RECOVERY_TOOLCALL_ORDER_FIX` (default off), mid-stream continuation becomes tool-call safe: any tool call seen in the stream — in flight or finished — blocks a continuation, and an empty continuation stops after one attempt. Maintainer rework before merge (kept the idea, no default behavior change): - The empty-continuation short-circuit also ran with the flag off; it is now gated, so the flag-off path uses the whole budget exactly as before (regression test added). - The latch re-arm that let a continuation fire after a completed `finish_reason: tool_calls` is gone; index-less tool calls on multi-choice payloads are now blocked too; ~150 lines of dead trace plumbing removed. Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 c0f92ec: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
# Conflicts: # docs/reference/FEATURE_FLAGS.md # open-sse/services/streamRecovery.ts # tests/unit/feature-flags-settings.test.ts # tests/unit/server-owned-tool-loop-flag.test.ts
diegosouzapw
merged commit Sep 15, 2026
931c9f9
into
diegosouzapw:release/v3.8.51
0 of 2 checks passed
This was referenced Sep 24, 2026
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…iegosouzapw#13650) Recovery traces for mid-stream continuation: one `onContinueOutcome` hook reports suffix stitched, overlap rejected, terminal, empty, no-stream and refused (with reason), logged through `chatCore` at debug level; warn is reserved for the cases where recovery gives up. Maintainer rework before merge (kept the idea, no default behavior change): - The original logged a warn-level latch line on every streamed tool call; nominal and tool-call streams are now silent, and the existing `mid-stream continuation attempt N/4` line keeps its format. - `chatCore.ts` ends 3 lines shorter than the tip, so the baseline bump the PR carried was removed; the wiring is tested through a real `handleChatCore` continuation. Validated first on the combined board of all 38 PRs of this batch (10 merged as-is, 28 after the maintainer rework) on top of release/v3.8.51 d61b804: typecheck:core, check:open-sse-typecheck and check:dashboard-typecheck clean; ESLint clean on every changed file; file-size (rebaselined for the combined growth), complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync, migration-numbering and i18n new-key gates green; 735 focused node:test cases with the only batch-caused failure (a flag-count assertion) fixed. Then re-validated alone on the fresh release tip right before this merge: ESLint on the changed files, typecheck:core, check:open-sse-typecheck, the file-size/complexity/changelog gates and this PR's own tests. Thanks @maxmad64bis!
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Continuation attempts were traced in code but never logged: the hook from #13633 fired into the void, so a wrongly resumed stream stayed undetectable in prod. Now each attempt logs one
warn("STREAM_RECOVERY", …)line keyed by itsattemptcounter, joinable withonContinue. Nominal requests stay silent (zero lines on clean[DONE]).Related Issues
onRecoveryTrace,StreamRecoveryTrace, 5 outcomes)Validation
npm run lintTests Added Or Updated
tests/unit/stream-recovery-trace-logging.test.ts: line format per kind (attempt/outcome/latch),attemptjoin withonContinue, zero lines on nominal done, flag-off silence — 9 tests, 9/9 green; neighbors 39/39 (continuation-wiring + toolcall)typecheck:corecleanCoverage Notes
chatCore.ts+2 lines at the existingonContinuesite; formatter extracted torecoveryTraceLogging.ts(24 lines). No decision touched (scan, latch, thresholds unchanged).Reviewer Notes
orderFix=appears onlatchlines only (the one transition the flag governs); attempt/outcome lines omit it so dashboards don't pin an ephemeral flag name.refusedReason: "terminal"is an internal sentinel (never emitted: the nominal-silence gate swallows it); documented in the type.