Skip to content

fix(backend): eval runner opts out of output-style and memory injection - #13414

Merged
diegosouzapw merged 3 commits into
diegosouzapw:release/v3.8.51from
KooshaPari:pr/13139-eval-opt-out
Sep 15, 2026
Merged

diegosouzapw merged 3 commits into
diegosouzapw:release/v3.8.51from
KooshaPari:pr/13139-eval-opt-out

Conversation

@KooshaPari

Copy link
Copy Markdown
Contributor

Summary

The eval runner now opts out of output-style and memory injection so that eval cases measure the model, not injected context.

Problem

executeEvalCase() sent cases down the ordinary chat path with only Content-Type and optionally Authorization headers. Three injections applied:

  1. Output styles: persona system messages prepended when a style is selected
  2. Memory context: stored text prepended as a system message when the request has a memory owner (API key)
  3. Memory + skills tools: memory_* tools appended to the request

These are documented opt-outs via x-omniroute-compression: off and x-omniroute-no-memory: true, but the eval runner set neither.

Effect: pass rates varied wildly depending on output-style selection and API key presence, measuring injected context rather than model capability.

Fix

Added both opt-out headers to the eval runner's request:

  • x-omniroute-compression: off — disables output-style persona injection
  • x-omniroute-no-memory: true — disables memory context and memory tools injection

Test results

tests 3 / pass 3 / fail 0

Fixes #13139

The 'does not duplicate custom Jina specialty models' test asserted
catalog IDs prefixed with 'jina-ai/' (the connection/provider ID), but
the catalog builder uses the provider alias 'jina' — which was
introduced in v3.8.36. The test above it in the same file ('does not
duplicate imported Jina specialty models') already asserts 'jina/' and
passes; this second test was simply never updated.

Fix the two assert.equal strings from 'jina-ai/' to 'jina/' to match
the current product behavior.

Fixes diegosouzapw#13313 (Jina portion — pt-BR compression integration tests need
separate investigation)
…iegosouzapw#13308)

The VACUUM INTO path in createManagedDbBackup() — reached from the
health-check-repair flow — never called the retention helper, so
db_backups/ grew without bound (observed: 10+ GB across 18 snapshots
in 8 days).

The sibling path (backupDbFile → backup.ts) already calls
cleanupDbBackups() after each snapshot; the health-check path simply
forgot to do the same.

Fix: after a successful VACUUM INTO in createManagedDbBackup(), call
pruneBackupDirectory() from backupRetention.ts using the same
env-var precedence (DB_BACKUP_MAX_FILES / DB_BACKUP_RETENTION_DAYS)
as the manual/scheduled backup path. Pruning is wrapped in a try/catch
so a best-effort failure never obscures the backup result.

Regression test: db-backup-healthcheck-prune-13308.test.ts seeds a
backup directory with MAX_DB_BACKUPS + 5 families and asserts that
pruneBackupDirectory removes exactly the overflow.
…on (diegosouzapw#13139)

The eval runner sent cases down the ordinary chat path, so every case
picked up output-style persona messages and memory context/tools.
There was no way to opt out, so evals measured injected context as
much as the model.

Added x-omniroute-compression: off and x-omniroute-no-memory: true
headers to the eval runner's request, matching the documented opt-out
mechanism used by self-managed clients.

Added 3 regression tests verifying the headers are set correctly.

Fixes diegosouzapw#13139
Copilot AI lite review requested due to automatic review settings September 12, 2026 07:05

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@diegosouzapw
diegosouzapw merged commit 76c928d into diegosouzapw:release/v3.8.51 Sep 15, 2026
8 of 16 checks passed
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…on (diegosouzapw#13414)

The eval runner sends `x-omniroute-compression: off` and `x-omniroute-no-memory: true`, so cases measure the model rather than injected output styles or retrieved memory (diegosouzapw#13139). Both headers are the existing per-request opt-outs honored by `chatCore` (`open-sse/handlers/chatCore/headers.ts`).

Validated in one consolidated batch of this series (37 PRs boarded together on `release/v3.8.51`): `typecheck:core`, `check:open-sse-typecheck` and `check:dashboard-typecheck` clean; ESLint clean on every changed file; file-size, complexity, cognitive-complexity, changelog-integrity, docs-counts, docs-sync and migration-numbering gates green (only the pre-existing `open-sse/utils/stream.ts` file-size red remains, inherited from the base); 3,743 focused `node:test` cases plus 34 vitest cases green.

Thanks @KooshaPari!
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(backend): Eval runner cannot opt out of output-style and memory injection — evals measure injected context, not the model

3 participants