Skip to content

fix(compression): match legacy output text in hu, ja and zh - #15591

Merged
diegosouzapw merged 3 commits into
diegosouzapw:release/v3.8.52from
woodsonl:fix/output-styles-language-drift
Oct 6, 2026
Merged

diegosouzapw merged 3 commits into
diegosouzapw:release/v3.8.52from
woodsonl:fix/output-styles-language-drift

Conversation

@woodsonl

@woodsonl woodsonl commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

What broke

Legacy cavemanOutputMode configs (the pre-styles compression toggle) reach the request pipeline through the back-compat shim, which maps them to the terse-prose output style. Its injected text matched the old caveman injector in 9 of the 12 legacy languages. Two languages did not:

  • Hungarian got the English text because the terse-prose i18n map had no hu entry, so the style fell back to CAVEMAN_INSTRUCTION_BY_LANGUAGE.en.
  • Japanese and Chinese got one extra space before the shared boundary sentence. The old injector rejoined the stripped instruction and the boundary block with a fixed space; every ja and zh text in the catalog runs straight into that sentence.

The fix

  • catalog.ts: terse-prose gains CAVEMAN_INSTRUCTION_BY_LANGUAGE.hu, and the i18n map is now derived from the legacy table by destructuring, so a language added there cannot be missed here again.
  • apply.ts: the boundary block follows the instruction body with the spacing the last style's own text puts before SHARED_BOUNDARIES: one space when that text has whitespace (or nothing) before the marker, none when a non-whitespace character directly precedes it. This also removes the extra space from multi-style ja/zh selections and from terse-cjk.

For existing ja/zh and terse-cjk selections the injected text loses that space, so prompt-cache prefixes for those tuples invalidate once on upgrade. The byte change is the intended fix.

Tests

  • tests/unit/compression/output-styles-legacy-parity.test.ts (new): every language in CAVEMAN_INSTRUCTION_BY_LANGUAGE at lite/full/ultra through both injectors, comparing the system message text below each marker line. On the unfixed code, 9 of its 36 cases fail (hu, ja, zh); all pass on this branch.
  • tests/unit/compression/output-styles-boundaries.test.ts: pins the character before the boundary sentence for multi-style and terse-cjk selections (none for ja/zh, one space for hu via the English fallback), now at full and ultra levels.
  • tests/unit/compression/output-styles-i18n-matrix.test.ts: adds hu to the terse-prose baseline.

Verification

  • Focused output-styles suite: 134/134 pass
  • Broader compression suite: 86 pass, 1 pre-existing hard-coded skip
  • npm run typecheck:core: clean
  • ESLint on all changed files: clean

Notes for reviewers

Legacy cavemanOutputMode configs reach upstream through
applyOutputStyles, which maps them to the terse-prose style. Its text
matched the old caveman injector in 9 of the 12 legacy languages.
Hungarian got the English text, because the terse-prose i18n map had
no hu entry. Japanese and Chinese got one extra space before the
shared boundary sentence: the injector rejoined the stripped
instruction and the boundary block with a fixed space, while every ja
and zh text in the catalog runs straight into that sentence.

- catalog.ts: terse-prose takes CAVEMAN_INSTRUCTION_BY_LANGUAGE.hu.
- apply.ts: the boundary block follows the instruction body with the
  spacing the last style's own text puts before SHARED_BOUNDARIES.
  This also removes the space from multi-style ja and zh selections
  and from terse-cjk.
- The catalog.ts, apply.ts and backCompat.ts comments now say the text
  below the marker line matches the legacy injection. The marker line
  itself still differs.

Tests:
- output-styles-legacy-parity.test.ts (new) runs every language in
  CAVEMAN_INSTRUCTION_BY_LANGUAGE at lite, full and ultra through both
  injectors and compares the system message text below each marker.
  On the unfixed code, 9 of its 36 cases fail (hu, ja and zh).
- output-styles-boundaries.test.ts checks the character before the
  boundary sentence in multi-style and terse-cjk selections: none for
  ja and zh, one space for hu, where the last style falls back to its
  English text.
- output-styles-i18n-matrix.test.ts adds hu to the terse-prose
  baseline.
…s with behavior

- catalog.ts derives terse-prose levels/i18n from CAVEMAN_INSTRUCTION_BY_LANGUAGE by
  destructuring instead of hand-copying 11 keys, so a language added to the legacy
  table cannot be missed the way hu was.
- apply.ts comments state the separator rule as the invariant it implements (any
  whitespace before SHARED_BOUNDARIES collapses to one space; a non-whitespace
  character suppresses it) instead of enumerating ja and zh.
- The legacy-parity helper reads the instruction from every text surface (top-level
  system field or trailing system message), so a placement change cannot fake a
  failure.
- The multi-style spacing matrix gains an ultra-level case; a stale regression-guard
  comment now points at the parity suite.
@woodsonl
woodsonl requested a review from diegosouzapw as a code owner October 5, 2026 21:42
@diegosouzapw
diegosouzapw merged commit 44efe61 into diegosouzapw:release/v3.8.52 Oct 6, 2026
44 of 51 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants