Skip to content

fix(build): co-locate llmlingua SLM optionals into dist/node_modules (postinstall) - #4286

Merged
diegosouzapw merged 1 commit into
release/v3.8.30from
fix/slm-optionals-colocate-postinstall
Jun 19, 2026
Merged

diegosouzapw merged 1 commit into
release/v3.8.30from
fix/slm-optionals-colocate-postinstall

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

Problem

The compression ultra SLM tier (#4257) runs @atjsh/llmlingua-2 + @huggingface/transformers + @tensorflow/tfjs + js-tiktoken inside a worker thread shipped under dist/ (open-sse/services/compression/engines/llmlingua/onnxWorker.js). These are optionalDependencies — npm installs them into the root node_modules on --include=optional, but the Next.js standalone trace bundles only @huggingface/transformers (3.5.2, pinned) into dist/node_modules; it does not trace the dynamically-imported optional SLM packages.

The instance-split bug

The worker lives under dist/, so its import("@huggingface/transformers") resolves dist/node_modules (3.5.2) and it sets the model cacheDir on that instance's env. But its import("@atjsh/llmlingua-2") walks past dist/node_modules up to the root node_modules, and llmlingua-2's own transformers import then resolves a different instance. The cacheDir/localModelPath the worker configured never reaches the instance llmlingua-2 actually uses → the local model under ${DATA_DIR}/models/llmlingua is never found → the SLM tier silently fails-open (no compression). If the root transformers is a 4.x line, llmlingua-2 also throws on a tokenizer-API change (decoder.decode undefined).

This was observed live on the production VPS: dist/node_modules had transformers 3.5.2 but no @atjsh, while the optionals sat in the root tree → the SLM tier produced zero compression.

Fix

scripts/build/colocateOptionals.mjs co-locates the SLM optional closure from the root node_modules into dist/node_modules (no-clobber, so the pinned dist transformers 3.5.2 / onnxruntime / sharp stay). Then the worker resolves @atjsh/llmlingua-2 and @huggingface/transformers from the same dist/node_modules — a single 3.5.2 instance — so the env config applies and the local model loads.

  • @huggingface/transformers is intentionally not a closure seed: it is a peer of @atjsh/llmlingua-2 (not a regular dependency) and already lives in dist/node_modules, so the closure walk never reaches it (and no-clobber would skip it anyway).
  • Wired into scripts/build/postinstall.mjs next to the existing ensureSwcHelpers (same "copy from root → dist/node_modules" pattern already used for better-sqlite3 / wreq-js / @swc/helpers).
  • Idempotent + fail-soft: no-op when the optionals are absent (the common case — they are optional) or already co-located; a per-package copy failure only disables the SLM tier, which is itself fail-open, so it never fails the install.

Validation (Hard Rule #18)

VPS live test — the exact co-located layout (cp -rn of the closure into dist/node_modules) produced real 54.8% compression (11520 → 5203 chars) via real ONNX inference on 192.168.0.15, for both the default and the #4257 modelPath code paths:

Path ok shrink latency
default (no modelPath) ✅ 54.8% 3.0s cold
#4257 modelPath ✅ 54.8% 2.3s warm

Unit — tests/unit/colocate-optionals.test.ts (6 tests): closure walk skips the transformers peer; no-clobber preserves dist's pinned 3.5.2; idempotence; both skip-gates. tests/unit/pack-artifact-policy.test.ts updated for the new required/allowed path.

Changes

  • scripts/build/colocateOptionals.mjs (new) — closure walk + no-clobber co-location
  • scripts/build/postinstall.mjs — call ensureLlmlinguaOptionals()
  • package.json files + scripts/build/pack-artifact-policy.ts (allow + required lists)
  • tests/unit/colocate-optionals.test.ts (new) + tests/unit/pack-artifact-policy.test.ts
  • docs/ops/RELEASE_CHECKLIST.md — note the auto co-location

🤖 Generated with Claude Code

…(postinstall)

The compression "ultra" SLM tier (#4257) runs @atjsh/llmlingua-2 + transformers + tfjs
+ js-tiktoken in a worker thread shipped under dist/. These are optionalDependencies
installed into the ROOT node_modules on --include=optional, but the Next.js standalone
trace bundles ONLY @huggingface/transformers (3.5.2, pinned) into dist/node_modules —
not the dynamically-imported optionals.

Result: the worker resolves transformers from dist/node_modules (3.5.2) for its env
config but resolves @atjsh/llmlingua-2 from the ROOT, whose own transformers import
hits a DIFFERENT instance. The cacheDir config never reaches the instance llmlingua-2
uses, so the local model never loads and the SLM tier silently fails-open (and on a
root transformers 4.x, llmlingua-2 throws on the tokenizer API change).

Fix: postinstall co-locates the SLM optional closure from the root node_modules into
dist/node_modules (no-clobber, so the pinned dist transformers/onnxruntime stay), so
the worker resolves a single 3.5.2 instance and the local model loads.

VPS-validated (Rule #18): the co-located layout produced real 54.8% compression
(11520->5203 chars) via real ONNX inference on the production host, both the default
and the #4257 modelPath code paths.

- scripts/build/colocateOptionals.mjs: closure walk (deps+optionalDeps, skips the
  transformers peer) + no-clobber co-location; idempotent + fail-soft
- wired into scripts/build/postinstall.mjs next to ensureSwcHelpers
- registered in package.json files + pack-artifact allow/required lists
- tests/unit/colocate-optionals.test.ts: closure, no-clobber, idempotence, gates
- docs/ops/RELEASE_CHECKLIST.md: note the auto co-location
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a new script, colocateOptionals.mjs, to co-locate the LLMLingua-2 optional dependency closure into dist/node_modules, resolving an instance-split bug with @huggingface/transformers in the standalone bundle. It also integrates this script into the postinstall process and adds comprehensive unit tests. The review feedback suggests two key improvements: dereferencing symlinks during the copy process to support package managers like pnpm, and verifying the presence of all seed packages in dist/node_modules to ensure a robust, complete co-location check.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

if (existsSync(dest)) continue; // no-clobber: keep dist's pinned copy (transformers 3.5.2, …)
try {
mkdirSync(dirname(dest), { recursive: true });
cpSync(join(rootNm, name), dest, { recursive: true });

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

When copying optional dependencies from the root node_modules to dist/node_modules, we should dereference symlinks. If the user is using a package manager like pnpm (which is supported in this repository), the root node_modules will contain symlinks pointing to the virtual store (.pnpm). Copying these symlinks without dereferencing will result in broken symlinks in dist/node_modules because the relative path to the .pnpm store will no longer resolve correctly from the dist/ subdirectory. Setting dereference: true ensures that the actual files are copied, making the standalone bundle fully self-contained.

Suggested change
cpSync(join(rootNm, name), dest, { recursive: true });
cpSync(join(rootNm, name), dest, { recursive: true, dereference: true });

Comment on lines +120 to +122
if (existsSync(join(distNm, "@atjsh", "llmlingua-2"))) {
return { skipped: true, reason: "already co-located" };
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Checking only the first package (@atjsh/llmlingua-2) to determine if the co-location has already run is not fully robust. If a previous installation or postinstall script was interrupted or failed halfway, some of the other seed packages (like @tensorflow/tfjs or js-tiktoken) might still be missing from dist/node_modules. Checking that all seed packages exist in dist/node_modules ensures completeness and prevents partial/broken co-locations on subsequent runs.

Suggested change
if (existsSync(join(distNm, "@atjsh", "llmlingua-2"))) {
return { skipped: true, reason: "already co-located" };
}
if (SEED_PACKAGES.every((seed) => existsSync(join(distNm, seed)))) {
return { skipped: true, reason: "already co-located" };
}

@diegosouzapw
diegosouzapw merged commit 0acb8d0 into release/v3.8.30 Jun 19, 2026
4 checks passed
diegosouzapw added a commit that referenced this pull request Jun 20, 2026
…on flow (#4323)

* fix(compression): SLM worker resolves deps+worker file without import.meta.url (B-SLM)

The Next.js standalone bundle (webpack) replaces createRequire(import.meta.url)
with a stub that always throws MODULE_NOT_FOUND, and freezes import.meta.url to the
build-machine path. So depsAvailable() was always false (the worker never spawned)
and resolveWorkerFile() anchored on a path absent at runtime — the SLM silently
fell back to the aggressive summarizer in production. Confirmed by inspecting
dist/.build/next/server/chunks/26410.js (stub module 215743 + frozen file:// path).

Replace both with filesystem probing from runtime anchors (process.cwd(),
process.argv[1]) that survive the bundle. Necessary complement to #4286 (deps
co-location) for the SLM to actually engage in prod; still fail-open without it.
VPS live validation deferred (Rule #18); local resolver regression tests added.

* fix(compression): ultra heuristic preserves code blocks / inline code / URLs (B-ULTRA-CODE)

ultra.ts called pruneByScore on raw text with no tombstoning, so the token pruner
dropped low-score code tokens (`b)`, `{`, `+`) inside fenced blocks while leaving the
fence markers intact — output that looked like valid code but was syntactically
destroyed. caveman + llmlingua both extract/restore preserved blocks first; ultra was
the only pruning engine that didn't.

Add pruneProseOnly(): extractPreservedBlocks tombstones fenced code, inline code,
URLs, CONST_CASE, versions; only the prose between placeholders is pruned; preserved
blocks are re-stitched verbatim.

* fix(compression): GCF round-trips values containing the inline-array pattern [..]: (B-GCF-QUOTE)

A value like `ERR[404]: Not Found` / `[Speaker 1]: Hello` nested one level deep was
emitted bare and re-parsed by the decoder as an inline-array header → it threw
`count_mismatch` (or silently decoded wrong), losing the whole block. headroomEngine
.apply() ships such blobs in prod, so this was a reachable lossless violation.

Two complementary fixes, both per SPEC §2.4:
- encode: needsQuote() now quotes strings matching `[`…`]``:` (spec compliance / other
  decoders).
- decode: the inline-array branch only fires when the bracket is in the KEY position
  (no `=` before it), so a quoted `note="ERR[404]: …"` value falls through to key=value.

* fix(compression): aggressive fidelity — keep text blocks, compress Anthropic tool_result, don't corrupt JSON (B-AGG-*)

Three fidelity fixes in the aggressive path (each TDD, aggressive-fidelity.test.ts):
- B-AGG-TEXTDROP: replaceTextContent dropped 2nd+ text blocks unconditionally; now a
  trailing block is dropped only when its text is already subsumed by newText, else kept.
- B-AGG-ANTHROPIC-TR: tool-result compression only fired for OpenAI role:tool messages;
  now Anthropic-shape tool_result content blocks (inside user messages) are compressed
  too, preserving tool_use_id + block structure.
- B-AGG-JSONTAG: the [COMPRESSED:aging:*] prefix corrupted JSON/code payloads; pure JSON
  is now kept verbatim+untagged (stays parseable), fenced blocks get the tag on a
  preceding line.

* fix(compression): accessibility collapse preserves [ref] anchors + fires on interleaved trees (B-MCPA11Y-*)

- B-MCPA11Y-ANCHORS: collapseRepeated silently dropped the omitted middle siblings'
  [ref=eNN] anchors (the agent could no longer click them); now every omitted ref is
  kept alongside the collapse notice. Wires the previously-dead preserveRefPattern.
  Invariant: extractRefs(input) ⊆ extractRefs(output).
- B-MCPA11Y-COLLAPSE: noise removal blanked lines (replace→""), and a blank line broke
  the sibling run so collapse never fired on realistic interleaved trees; noise lines
  are now deleted, and the sibling walk skips stray blanks.

* fix(compression): rtk intensity scales the line budget (B-RTK-INTENSITY)

The intensity knob only set smartTruncate's preserveHead/Tail (16↔24), which rarely
fired because the matched filter capped lines first — so minimal/standard/aggressive
produced byte-identical output on filter-matched tool output. effectiveMaxLines() now
scales the effective line budget (minimal 1.5x, standard 1x, aggressive 0.5x) at both
the per-filter and engine-level truncation sites. Both go through smartTruncate with
priorityPatterns, so error/failure lines survive at every intensity (tested).

* fix(compression): robust language detection + auto-detect honors the detected pack (B-LANG-*)

- B-LANG-DETECTOR: detector was first-match-wins on a single keyword, and some hints are
  English-ambiguous ("configuration" in fr, "error" in es) → English text misclassified.
  Now score-based (count native-keyword hits, highest wins), and the two English-ambiguous
  words are removed from the hint lists, so a lone shared word never misclassifies while
  sparse-keyword languages (id) still detect on a single native word.
- B-LANG-DORMANT: with autoDetectLanguage on but enabledPacks ["en"], detected non-English
  text fell back to the English pack, whose `articles` rule deletes foreign articles
  (pt-BR "a"/"o"). Auto-detect now uses the detected pack directly (it always has rules);
  enabledPacks still gates manual selection.

* fix(compression): mode selection enables its engine + align stacked allowlist (B-MODE-ENGINE-DECOUPLE, B-PIPELINE-DIVERGENCE)

- B-MODE-ENGINE-DECOUPLE: picking the standard/rtk MODE now runs caveman/rtk regardless of
  the per-engine enabled flag — the mode selection is the enable signal (the per-engine flag
  still gates stacked pipeline steps). Previously an operator who picked a mode but left the
  engine toggle off got silent 0% compression.
- B-PIPELINE-DIVERGENCE: the global stackedPipeline normalizer stripped
  session-dedup/ccr/headroom/llmlingua (engines the combo path accepts via KNOWN_ENGINE_IDS).
  The allowlist now matches, so the global setting can use all registered engines.

* docs(compression): correct SLM "stable" claim + document partial packs / stacked telemetry limits

- The llmlingua `stable:true` comment claimed the bundle walk-up + deps-gate were
  "confirmed against the live install" — that was wrong (webpack froze import.meta.url and
  stubbed createRequire, so the worker never spawned in prod). Corrected to reflect B-SLM.
- COMPRESSION_ENGINES.md: add a Known limitations section (SLM dep co-location requirement,
  partial de/fr/ja packs, no-op engines absent from engineBreakdown).

* fix(compression): cast normalized engine id to CompressionPipelineStep['engine'] (typecheck)
@diegosouzapw diegosouzapw mentioned this pull request Jun 20, 2026
@diegosouzapw
diegosouzapw deleted the fix/slm-optionals-colocate-postinstall branch June 21, 2026 12:33
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
…(postinstall) (diegosouzapw#4286)

The compression "ultra" SLM tier (diegosouzapw#4257) runs @atjsh/llmlingua-2 + transformers + tfjs
+ js-tiktoken in a worker thread shipped under dist/. These are optionalDependencies
installed into the ROOT node_modules on --include=optional, but the Next.js standalone
trace bundles ONLY @huggingface/transformers (3.5.2, pinned) into dist/node_modules —
not the dynamically-imported optionals.

Result: the worker resolves transformers from dist/node_modules (3.5.2) for its env
config but resolves @atjsh/llmlingua-2 from the ROOT, whose own transformers import
hits a DIFFERENT instance. The cacheDir config never reaches the instance llmlingua-2
uses, so the local model never loads and the SLM tier silently fails-open (and on a
root transformers 4.x, llmlingua-2 throws on the tokenizer API change).

Fix: postinstall co-locates the SLM optional closure from the root node_modules into
dist/node_modules (no-clobber, so the pinned dist transformers/onnxruntime stay), so
the worker resolves a single 3.5.2 instance and the local model loads.

VPS-validated (Rule diegosouzapw#18): the co-located layout produced real 54.8% compression
(11520->5203 chars) via real ONNX inference on the production host, both the default
and the diegosouzapw#4257 modelPath code paths.

- scripts/build/colocateOptionals.mjs: closure walk (deps+optionalDeps, skips the
  transformers peer) + no-clobber co-location; idempotent + fail-soft
- wired into scripts/build/postinstall.mjs next to ensureSwcHelpers
- registered in package.json files + pack-artifact allow/required lists
- tests/unit/colocate-optionals.test.ts: closure, no-clobber, idempotence, gates
- docs/ops/RELEASE_CHECKLIST.md: note the auto co-location
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
…on flow (diegosouzapw#4323)

* fix(compression): SLM worker resolves deps+worker file without import.meta.url (B-SLM)

The Next.js standalone bundle (webpack) replaces createRequire(import.meta.url)
with a stub that always throws MODULE_NOT_FOUND, and freezes import.meta.url to the
build-machine path. So depsAvailable() was always false (the worker never spawned)
and resolveWorkerFile() anchored on a path absent at runtime — the SLM silently
fell back to the aggressive summarizer in production. Confirmed by inspecting
dist/.build/next/server/chunks/26410.js (stub module 215743 + frozen file:// path).

Replace both with filesystem probing from runtime anchors (process.cwd(),
process.argv[1]) that survive the bundle. Necessary complement to diegosouzapw#4286 (deps
co-location) for the SLM to actually engage in prod; still fail-open without it.
VPS live validation deferred (Rule diegosouzapw#18); local resolver regression tests added.

* fix(compression): ultra heuristic preserves code blocks / inline code / URLs (B-ULTRA-CODE)

ultra.ts called pruneByScore on raw text with no tombstoning, so the token pruner
dropped low-score code tokens (`b)`, `{`, `+`) inside fenced blocks while leaving the
fence markers intact — output that looked like valid code but was syntactically
destroyed. caveman + llmlingua both extract/restore preserved blocks first; ultra was
the only pruning engine that didn't.

Add pruneProseOnly(): extractPreservedBlocks tombstones fenced code, inline code,
URLs, CONST_CASE, versions; only the prose between placeholders is pruned; preserved
blocks are re-stitched verbatim.

* fix(compression): GCF round-trips values containing the inline-array pattern [..]: (B-GCF-QUOTE)

A value like `ERR[404]: Not Found` / `[Speaker 1]: Hello` nested one level deep was
emitted bare and re-parsed by the decoder as an inline-array header → it threw
`count_mismatch` (or silently decoded wrong), losing the whole block. headroomEngine
.apply() ships such blobs in prod, so this was a reachable lossless violation.

Two complementary fixes, both per SPEC §2.4:
- encode: needsQuote() now quotes strings matching `[`…`]``:` (spec compliance / other
  decoders).
- decode: the inline-array branch only fires when the bracket is in the KEY position
  (no `=` before it), so a quoted `note="ERR[404]: …"` value falls through to key=value.

* fix(compression): aggressive fidelity — keep text blocks, compress Anthropic tool_result, don't corrupt JSON (B-AGG-*)

Three fidelity fixes in the aggressive path (each TDD, aggressive-fidelity.test.ts):
- B-AGG-TEXTDROP: replaceTextContent dropped 2nd+ text blocks unconditionally; now a
  trailing block is dropped only when its text is already subsumed by newText, else kept.
- B-AGG-ANTHROPIC-TR: tool-result compression only fired for OpenAI role:tool messages;
  now Anthropic-shape tool_result content blocks (inside user messages) are compressed
  too, preserving tool_use_id + block structure.
- B-AGG-JSONTAG: the [COMPRESSED:aging:*] prefix corrupted JSON/code payloads; pure JSON
  is now kept verbatim+untagged (stays parseable), fenced blocks get the tag on a
  preceding line.

* fix(compression): accessibility collapse preserves [ref] anchors + fires on interleaved trees (B-MCPA11Y-*)

- B-MCPA11Y-ANCHORS: collapseRepeated silently dropped the omitted middle siblings'
  [ref=eNN] anchors (the agent could no longer click them); now every omitted ref is
  kept alongside the collapse notice. Wires the previously-dead preserveRefPattern.
  Invariant: extractRefs(input) ⊆ extractRefs(output).
- B-MCPA11Y-COLLAPSE: noise removal blanked lines (replace→""), and a blank line broke
  the sibling run so collapse never fired on realistic interleaved trees; noise lines
  are now deleted, and the sibling walk skips stray blanks.

* fix(compression): rtk intensity scales the line budget (B-RTK-INTENSITY)

The intensity knob only set smartTruncate's preserveHead/Tail (16↔24), which rarely
fired because the matched filter capped lines first — so minimal/standard/aggressive
produced byte-identical output on filter-matched tool output. effectiveMaxLines() now
scales the effective line budget (minimal 1.5x, standard 1x, aggressive 0.5x) at both
the per-filter and engine-level truncation sites. Both go through smartTruncate with
priorityPatterns, so error/failure lines survive at every intensity (tested).

* fix(compression): robust language detection + auto-detect honors the detected pack (B-LANG-*)

- B-LANG-DETECTOR: detector was first-match-wins on a single keyword, and some hints are
  English-ambiguous ("configuration" in fr, "error" in es) → English text misclassified.
  Now score-based (count native-keyword hits, highest wins), and the two English-ambiguous
  words are removed from the hint lists, so a lone shared word never misclassifies while
  sparse-keyword languages (id) still detect on a single native word.
- B-LANG-DORMANT: with autoDetectLanguage on but enabledPacks ["en"], detected non-English
  text fell back to the English pack, whose `articles` rule deletes foreign articles
  (pt-BR "a"/"o"). Auto-detect now uses the detected pack directly (it always has rules);
  enabledPacks still gates manual selection.

* fix(compression): mode selection enables its engine + align stacked allowlist (B-MODE-ENGINE-DECOUPLE, B-PIPELINE-DIVERGENCE)

- B-MODE-ENGINE-DECOUPLE: picking the standard/rtk MODE now runs caveman/rtk regardless of
  the per-engine enabled flag — the mode selection is the enable signal (the per-engine flag
  still gates stacked pipeline steps). Previously an operator who picked a mode but left the
  engine toggle off got silent 0% compression.
- B-PIPELINE-DIVERGENCE: the global stackedPipeline normalizer stripped
  session-dedup/ccr/headroom/llmlingua (engines the combo path accepts via KNOWN_ENGINE_IDS).
  The allowlist now matches, so the global setting can use all registered engines.

* docs(compression): correct SLM "stable" claim + document partial packs / stacked telemetry limits

- The llmlingua `stable:true` comment claimed the bundle walk-up + deps-gate were
  "confirmed against the live install" — that was wrong (webpack froze import.meta.url and
  stubbed createRequire, so the worker never spawned in prod). Corrected to reflect B-SLM.
- COMPRESSION_ENGINES.md: add a Known limitations section (SLM dep co-location requirement,
  partial de/fr/ja packs, no-op engines absent from engineBreakdown).

* fix(compression): cast normalized engine id to CompressionPipelineStep['engine'] (typecheck)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant