Skip to content

fix(compression): use Unicode-aware word boundaries in Russian rule packs (#15677) - #15719

Merged
diegosouzapw merged 2 commits into
release/v3.8.52from
fix/15677-ru-compression-boundary
Oct 8, 2026
Merged

diegosouzapw merged 2 commits into
release/v3.8.52from
fix/15677-ru-compression-boundary

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

Closes #15677

Root cause

JS \b is ASCII-only and the Russian packs compiled with flags gi (no u), so 26 of 32 rules in open-sse/services/compression/rules/ru/*.json could never match Cyrillic (tokensSaved ~0). The detector's ru keyword regex had the same dead \b (masked by a catch-all Cyrillic hint).

Fix

Replaced \b with (?<![\p{L}\p{N}_]) / (?![\p{L}\p{N}_]) and set giu on affected ru rules (dedup \w+ -> [\p{L}\p{N}_]+); same fix for the ru detector keyword hint. Patterns remain bounded (no new quantifier nesting).

Not included (follow-up): user-level pack dir ~/.omniroute/compression/rules lookup in ruleLoader (docs mismatch), and the same \b issue in es/pt-BR/it/fr/de packs.

Tests

tests/unit/compression/ru-cyrillic-word-boundary-15677.test.ts (new). RED on unfixed code per triage probe (pleasantries rule left "Конечно, вот ответ." unchanged; detector keyword regex false); GREEN now (5/5). 12 related compression/detector test files: 104/104 pass.

Gates

file-size, complexity, cognitive-complexity, eslint, typecheck:core, compression-budget, pre-commit hooks.

⚠️ base-red inherited: #15306 — unit/integration/package-artifact gates timing out, vitest open-sse/mcp-server/tests/audit.test.ts, check:ai-attribution range error

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

CI Coverage Report

  • Coverage job: skipped
  • PR test policy: success

Coverage artifact was not available for this run.

…ru-compression-boundary

# Conflicts:
#	open-sse/services/compression/languageDetector.ts
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(backend): Russian language detector uses \b which never matches Cyrillic (ASCII-only in JavaScript)

1 participant