fix(build): co-locate llmlingua SLM optionals into dist/node_modules (postinstall) - #4286
Conversation
…(postinstall) The compression "ultra" SLM tier (#4257) runs @atjsh/llmlingua-2 + transformers + tfjs + js-tiktoken in a worker thread shipped under dist/. These are optionalDependencies installed into the ROOT node_modules on --include=optional, but the Next.js standalone trace bundles ONLY @huggingface/transformers (3.5.2, pinned) into dist/node_modules — not the dynamically-imported optionals. Result: the worker resolves transformers from dist/node_modules (3.5.2) for its env config but resolves @atjsh/llmlingua-2 from the ROOT, whose own transformers import hits a DIFFERENT instance. The cacheDir config never reaches the instance llmlingua-2 uses, so the local model never loads and the SLM tier silently fails-open (and on a root transformers 4.x, llmlingua-2 throws on the tokenizer API change). Fix: postinstall co-locates the SLM optional closure from the root node_modules into dist/node_modules (no-clobber, so the pinned dist transformers/onnxruntime stay), so the worker resolves a single 3.5.2 instance and the local model loads. VPS-validated (Rule #18): the co-located layout produced real 54.8% compression (11520->5203 chars) via real ONNX inference on the production host, both the default and the #4257 modelPath code paths. - scripts/build/colocateOptionals.mjs: closure walk (deps+optionalDeps, skips the transformers peer) + no-clobber co-location; idempotent + fail-soft - wired into scripts/build/postinstall.mjs next to ensureSwcHelpers - registered in package.json files + pack-artifact allow/required lists - tests/unit/colocate-optionals.test.ts: closure, no-clobber, idempotence, gates - docs/ops/RELEASE_CHECKLIST.md: note the auto co-location
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
There was a problem hiding this comment.
Code Review
This pull request introduces a new script, colocateOptionals.mjs, to co-locate the LLMLingua-2 optional dependency closure into dist/node_modules, resolving an instance-split bug with @huggingface/transformers in the standalone bundle. It also integrates this script into the postinstall process and adds comprehensive unit tests. The review feedback suggests two key improvements: dereferencing symlinks during the copy process to support package managers like pnpm, and verifying the presence of all seed packages in dist/node_modules to ensure a robust, complete co-location check.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| if (existsSync(dest)) continue; // no-clobber: keep dist's pinned copy (transformers 3.5.2, …) | ||
| try { | ||
| mkdirSync(dirname(dest), { recursive: true }); | ||
| cpSync(join(rootNm, name), dest, { recursive: true }); |
There was a problem hiding this comment.
When copying optional dependencies from the root node_modules to dist/node_modules, we should dereference symlinks. If the user is using a package manager like pnpm (which is supported in this repository), the root node_modules will contain symlinks pointing to the virtual store (.pnpm). Copying these symlinks without dereferencing will result in broken symlinks in dist/node_modules because the relative path to the .pnpm store will no longer resolve correctly from the dist/ subdirectory. Setting dereference: true ensures that the actual files are copied, making the standalone bundle fully self-contained.
| cpSync(join(rootNm, name), dest, { recursive: true }); | |
| cpSync(join(rootNm, name), dest, { recursive: true, dereference: true }); |
| if (existsSync(join(distNm, "@atjsh", "llmlingua-2"))) { | ||
| return { skipped: true, reason: "already co-located" }; | ||
| } |
There was a problem hiding this comment.
Checking only the first package (@atjsh/llmlingua-2) to determine if the co-location has already run is not fully robust. If a previous installation or postinstall script was interrupted or failed halfway, some of the other seed packages (like @tensorflow/tfjs or js-tiktoken) might still be missing from dist/node_modules. Checking that all seed packages exist in dist/node_modules ensures completeness and prevents partial/broken co-locations on subsequent runs.
| if (existsSync(join(distNm, "@atjsh", "llmlingua-2"))) { | |
| return { skipped: true, reason: "already co-located" }; | |
| } | |
| if (SEED_PACKAGES.every((seed) => existsSync(join(distNm, seed)))) { | |
| return { skipped: true, reason: "already co-located" }; | |
| } |
…on flow (#4323) * fix(compression): SLM worker resolves deps+worker file without import.meta.url (B-SLM) The Next.js standalone bundle (webpack) replaces createRequire(import.meta.url) with a stub that always throws MODULE_NOT_FOUND, and freezes import.meta.url to the build-machine path. So depsAvailable() was always false (the worker never spawned) and resolveWorkerFile() anchored on a path absent at runtime — the SLM silently fell back to the aggressive summarizer in production. Confirmed by inspecting dist/.build/next/server/chunks/26410.js (stub module 215743 + frozen file:// path). Replace both with filesystem probing from runtime anchors (process.cwd(), process.argv[1]) that survive the bundle. Necessary complement to #4286 (deps co-location) for the SLM to actually engage in prod; still fail-open without it. VPS live validation deferred (Rule #18); local resolver regression tests added. * fix(compression): ultra heuristic preserves code blocks / inline code / URLs (B-ULTRA-CODE) ultra.ts called pruneByScore on raw text with no tombstoning, so the token pruner dropped low-score code tokens (`b)`, `{`, `+`) inside fenced blocks while leaving the fence markers intact — output that looked like valid code but was syntactically destroyed. caveman + llmlingua both extract/restore preserved blocks first; ultra was the only pruning engine that didn't. Add pruneProseOnly(): extractPreservedBlocks tombstones fenced code, inline code, URLs, CONST_CASE, versions; only the prose between placeholders is pruned; preserved blocks are re-stitched verbatim. * fix(compression): GCF round-trips values containing the inline-array pattern [..]: (B-GCF-QUOTE) A value like `ERR[404]: Not Found` / `[Speaker 1]: Hello` nested one level deep was emitted bare and re-parsed by the decoder as an inline-array header → it threw `count_mismatch` (or silently decoded wrong), losing the whole block. headroomEngine .apply() ships such blobs in prod, so this was a reachable lossless violation. Two complementary fixes, both per SPEC §2.4: - encode: needsQuote() now quotes strings matching `[`…`]``:` (spec compliance / other decoders). - decode: the inline-array branch only fires when the bracket is in the KEY position (no `=` before it), so a quoted `note="ERR[404]: …"` value falls through to key=value. * fix(compression): aggressive fidelity — keep text blocks, compress Anthropic tool_result, don't corrupt JSON (B-AGG-*) Three fidelity fixes in the aggressive path (each TDD, aggressive-fidelity.test.ts): - B-AGG-TEXTDROP: replaceTextContent dropped 2nd+ text blocks unconditionally; now a trailing block is dropped only when its text is already subsumed by newText, else kept. - B-AGG-ANTHROPIC-TR: tool-result compression only fired for OpenAI role:tool messages; now Anthropic-shape tool_result content blocks (inside user messages) are compressed too, preserving tool_use_id + block structure. - B-AGG-JSONTAG: the [COMPRESSED:aging:*] prefix corrupted JSON/code payloads; pure JSON is now kept verbatim+untagged (stays parseable), fenced blocks get the tag on a preceding line. * fix(compression): accessibility collapse preserves [ref] anchors + fires on interleaved trees (B-MCPA11Y-*) - B-MCPA11Y-ANCHORS: collapseRepeated silently dropped the omitted middle siblings' [ref=eNN] anchors (the agent could no longer click them); now every omitted ref is kept alongside the collapse notice. Wires the previously-dead preserveRefPattern. Invariant: extractRefs(input) ⊆ extractRefs(output). - B-MCPA11Y-COLLAPSE: noise removal blanked lines (replace→""), and a blank line broke the sibling run so collapse never fired on realistic interleaved trees; noise lines are now deleted, and the sibling walk skips stray blanks. * fix(compression): rtk intensity scales the line budget (B-RTK-INTENSITY) The intensity knob only set smartTruncate's preserveHead/Tail (16↔24), which rarely fired because the matched filter capped lines first — so minimal/standard/aggressive produced byte-identical output on filter-matched tool output. effectiveMaxLines() now scales the effective line budget (minimal 1.5x, standard 1x, aggressive 0.5x) at both the per-filter and engine-level truncation sites. Both go through smartTruncate with priorityPatterns, so error/failure lines survive at every intensity (tested). * fix(compression): robust language detection + auto-detect honors the detected pack (B-LANG-*) - B-LANG-DETECTOR: detector was first-match-wins on a single keyword, and some hints are English-ambiguous ("configuration" in fr, "error" in es) → English text misclassified. Now score-based (count native-keyword hits, highest wins), and the two English-ambiguous words are removed from the hint lists, so a lone shared word never misclassifies while sparse-keyword languages (id) still detect on a single native word. - B-LANG-DORMANT: with autoDetectLanguage on but enabledPacks ["en"], detected non-English text fell back to the English pack, whose `articles` rule deletes foreign articles (pt-BR "a"/"o"). Auto-detect now uses the detected pack directly (it always has rules); enabledPacks still gates manual selection. * fix(compression): mode selection enables its engine + align stacked allowlist (B-MODE-ENGINE-DECOUPLE, B-PIPELINE-DIVERGENCE) - B-MODE-ENGINE-DECOUPLE: picking the standard/rtk MODE now runs caveman/rtk regardless of the per-engine enabled flag — the mode selection is the enable signal (the per-engine flag still gates stacked pipeline steps). Previously an operator who picked a mode but left the engine toggle off got silent 0% compression. - B-PIPELINE-DIVERGENCE: the global stackedPipeline normalizer stripped session-dedup/ccr/headroom/llmlingua (engines the combo path accepts via KNOWN_ENGINE_IDS). The allowlist now matches, so the global setting can use all registered engines. * docs(compression): correct SLM "stable" claim + document partial packs / stacked telemetry limits - The llmlingua `stable:true` comment claimed the bundle walk-up + deps-gate were "confirmed against the live install" — that was wrong (webpack froze import.meta.url and stubbed createRequire, so the worker never spawned in prod). Corrected to reflect B-SLM. - COMPRESSION_ENGINES.md: add a Known limitations section (SLM dep co-location requirement, partial de/fr/ja packs, no-op engines absent from engineBreakdown). * fix(compression): cast normalized engine id to CompressionPipelineStep['engine'] (typecheck)
…(postinstall) (diegosouzapw#4286) The compression "ultra" SLM tier (diegosouzapw#4257) runs @atjsh/llmlingua-2 + transformers + tfjs + js-tiktoken in a worker thread shipped under dist/. These are optionalDependencies installed into the ROOT node_modules on --include=optional, but the Next.js standalone trace bundles ONLY @huggingface/transformers (3.5.2, pinned) into dist/node_modules — not the dynamically-imported optionals. Result: the worker resolves transformers from dist/node_modules (3.5.2) for its env config but resolves @atjsh/llmlingua-2 from the ROOT, whose own transformers import hits a DIFFERENT instance. The cacheDir config never reaches the instance llmlingua-2 uses, so the local model never loads and the SLM tier silently fails-open (and on a root transformers 4.x, llmlingua-2 throws on the tokenizer API change). Fix: postinstall co-locates the SLM optional closure from the root node_modules into dist/node_modules (no-clobber, so the pinned dist transformers/onnxruntime stay), so the worker resolves a single 3.5.2 instance and the local model loads. VPS-validated (Rule diegosouzapw#18): the co-located layout produced real 54.8% compression (11520->5203 chars) via real ONNX inference on the production host, both the default and the diegosouzapw#4257 modelPath code paths. - scripts/build/colocateOptionals.mjs: closure walk (deps+optionalDeps, skips the transformers peer) + no-clobber co-location; idempotent + fail-soft - wired into scripts/build/postinstall.mjs next to ensureSwcHelpers - registered in package.json files + pack-artifact allow/required lists - tests/unit/colocate-optionals.test.ts: closure, no-clobber, idempotence, gates - docs/ops/RELEASE_CHECKLIST.md: note the auto co-location
…on flow (diegosouzapw#4323) * fix(compression): SLM worker resolves deps+worker file without import.meta.url (B-SLM) The Next.js standalone bundle (webpack) replaces createRequire(import.meta.url) with a stub that always throws MODULE_NOT_FOUND, and freezes import.meta.url to the build-machine path. So depsAvailable() was always false (the worker never spawned) and resolveWorkerFile() anchored on a path absent at runtime — the SLM silently fell back to the aggressive summarizer in production. Confirmed by inspecting dist/.build/next/server/chunks/26410.js (stub module 215743 + frozen file:// path). Replace both with filesystem probing from runtime anchors (process.cwd(), process.argv[1]) that survive the bundle. Necessary complement to diegosouzapw#4286 (deps co-location) for the SLM to actually engage in prod; still fail-open without it. VPS live validation deferred (Rule diegosouzapw#18); local resolver regression tests added. * fix(compression): ultra heuristic preserves code blocks / inline code / URLs (B-ULTRA-CODE) ultra.ts called pruneByScore on raw text with no tombstoning, so the token pruner dropped low-score code tokens (`b)`, `{`, `+`) inside fenced blocks while leaving the fence markers intact — output that looked like valid code but was syntactically destroyed. caveman + llmlingua both extract/restore preserved blocks first; ultra was the only pruning engine that didn't. Add pruneProseOnly(): extractPreservedBlocks tombstones fenced code, inline code, URLs, CONST_CASE, versions; only the prose between placeholders is pruned; preserved blocks are re-stitched verbatim. * fix(compression): GCF round-trips values containing the inline-array pattern [..]: (B-GCF-QUOTE) A value like `ERR[404]: Not Found` / `[Speaker 1]: Hello` nested one level deep was emitted bare and re-parsed by the decoder as an inline-array header → it threw `count_mismatch` (or silently decoded wrong), losing the whole block. headroomEngine .apply() ships such blobs in prod, so this was a reachable lossless violation. Two complementary fixes, both per SPEC §2.4: - encode: needsQuote() now quotes strings matching `[`…`]``:` (spec compliance / other decoders). - decode: the inline-array branch only fires when the bracket is in the KEY position (no `=` before it), so a quoted `note="ERR[404]: …"` value falls through to key=value. * fix(compression): aggressive fidelity — keep text blocks, compress Anthropic tool_result, don't corrupt JSON (B-AGG-*) Three fidelity fixes in the aggressive path (each TDD, aggressive-fidelity.test.ts): - B-AGG-TEXTDROP: replaceTextContent dropped 2nd+ text blocks unconditionally; now a trailing block is dropped only when its text is already subsumed by newText, else kept. - B-AGG-ANTHROPIC-TR: tool-result compression only fired for OpenAI role:tool messages; now Anthropic-shape tool_result content blocks (inside user messages) are compressed too, preserving tool_use_id + block structure. - B-AGG-JSONTAG: the [COMPRESSED:aging:*] prefix corrupted JSON/code payloads; pure JSON is now kept verbatim+untagged (stays parseable), fenced blocks get the tag on a preceding line. * fix(compression): accessibility collapse preserves [ref] anchors + fires on interleaved trees (B-MCPA11Y-*) - B-MCPA11Y-ANCHORS: collapseRepeated silently dropped the omitted middle siblings' [ref=eNN] anchors (the agent could no longer click them); now every omitted ref is kept alongside the collapse notice. Wires the previously-dead preserveRefPattern. Invariant: extractRefs(input) ⊆ extractRefs(output). - B-MCPA11Y-COLLAPSE: noise removal blanked lines (replace→""), and a blank line broke the sibling run so collapse never fired on realistic interleaved trees; noise lines are now deleted, and the sibling walk skips stray blanks. * fix(compression): rtk intensity scales the line budget (B-RTK-INTENSITY) The intensity knob only set smartTruncate's preserveHead/Tail (16↔24), which rarely fired because the matched filter capped lines first — so minimal/standard/aggressive produced byte-identical output on filter-matched tool output. effectiveMaxLines() now scales the effective line budget (minimal 1.5x, standard 1x, aggressive 0.5x) at both the per-filter and engine-level truncation sites. Both go through smartTruncate with priorityPatterns, so error/failure lines survive at every intensity (tested). * fix(compression): robust language detection + auto-detect honors the detected pack (B-LANG-*) - B-LANG-DETECTOR: detector was first-match-wins on a single keyword, and some hints are English-ambiguous ("configuration" in fr, "error" in es) → English text misclassified. Now score-based (count native-keyword hits, highest wins), and the two English-ambiguous words are removed from the hint lists, so a lone shared word never misclassifies while sparse-keyword languages (id) still detect on a single native word. - B-LANG-DORMANT: with autoDetectLanguage on but enabledPacks ["en"], detected non-English text fell back to the English pack, whose `articles` rule deletes foreign articles (pt-BR "a"/"o"). Auto-detect now uses the detected pack directly (it always has rules); enabledPacks still gates manual selection. * fix(compression): mode selection enables its engine + align stacked allowlist (B-MODE-ENGINE-DECOUPLE, B-PIPELINE-DIVERGENCE) - B-MODE-ENGINE-DECOUPLE: picking the standard/rtk MODE now runs caveman/rtk regardless of the per-engine enabled flag — the mode selection is the enable signal (the per-engine flag still gates stacked pipeline steps). Previously an operator who picked a mode but left the engine toggle off got silent 0% compression. - B-PIPELINE-DIVERGENCE: the global stackedPipeline normalizer stripped session-dedup/ccr/headroom/llmlingua (engines the combo path accepts via KNOWN_ENGINE_IDS). The allowlist now matches, so the global setting can use all registered engines. * docs(compression): correct SLM "stable" claim + document partial packs / stacked telemetry limits - The llmlingua `stable:true` comment claimed the bundle walk-up + deps-gate were "confirmed against the live install" — that was wrong (webpack froze import.meta.url and stubbed createRequire, so the worker never spawned in prod). Corrected to reflect B-SLM. - COMPRESSION_ENGINES.md: add a Known limitations section (SLM dep co-location requirement, partial de/fr/ja packs, no-op engines absent from engineBreakdown). * fix(compression): cast normalized engine id to CompressionPipelineStep['engine'] (typecheck)
Problem
The compression ultra SLM tier (#4257) runs
@atjsh/llmlingua-2+@huggingface/transformers+@tensorflow/tfjs+js-tiktokeninside a worker thread shipped underdist/(open-sse/services/compression/engines/llmlingua/onnxWorker.js). These areoptionalDependencies— npm installs them into the rootnode_moduleson--include=optional, but the Next.js standalone trace bundles only@huggingface/transformers(3.5.2, pinned) intodist/node_modules; it does not trace the dynamically-imported optional SLM packages.The instance-split bug
The worker lives under
dist/, so itsimport("@huggingface/transformers")resolvesdist/node_modules(3.5.2) and it sets the modelcacheDiron that instance'senv. But itsimport("@atjsh/llmlingua-2")walks pastdist/node_modulesup to the rootnode_modules, and llmlingua-2's own transformers import then resolves a different instance. ThecacheDir/localModelPaththe worker configured never reaches the instance llmlingua-2 actually uses → the local model under${DATA_DIR}/models/llmlinguais never found → the SLM tier silently fails-open (no compression). If the root transformers is a 4.x line, llmlingua-2 also throws on a tokenizer-API change (decoder.decodeundefined).This was observed live on the production VPS:
dist/node_moduleshad transformers 3.5.2 but no@atjsh, while the optionals sat in the root tree → the SLM tier produced zero compression.Fix
scripts/build/colocateOptionals.mjsco-locates the SLM optional closure from the rootnode_modulesintodist/node_modules(no-clobber, so the pinneddisttransformers 3.5.2 / onnxruntime / sharp stay). Then the worker resolves@atjsh/llmlingua-2and@huggingface/transformersfrom the samedist/node_modules— a single 3.5.2 instance — so the env config applies and the local model loads.@huggingface/transformersis intentionally not a closure seed: it is a peer of@atjsh/llmlingua-2(not a regular dependency) and already lives indist/node_modules, so the closure walk never reaches it (and no-clobber would skip it anyway).scripts/build/postinstall.mjsnext to the existingensureSwcHelpers(same "copy from root → dist/node_modules" pattern already used for better-sqlite3 / wreq-js / @swc/helpers).Validation (Hard Rule #18)
VPS live test — the exact co-located layout (
cp -rnof the closure intodist/node_modules) produced real 54.8% compression (11520 → 5203 chars) via real ONNX inference on192.168.0.15, for both the default and the #4257modelPathcode paths:Unit —
tests/unit/colocate-optionals.test.ts(6 tests): closure walk skips the transformers peer; no-clobber preserves dist's pinned 3.5.2; idempotence; both skip-gates.tests/unit/pack-artifact-policy.test.tsupdated for the new required/allowed path.Changes
scripts/build/colocateOptionals.mjs(new) — closure walk + no-clobber co-locationscripts/build/postinstall.mjs— callensureLlmlinguaOptionals()package.jsonfiles+scripts/build/pack-artifact-policy.ts(allow + required lists)tests/unit/colocate-optionals.test.ts(new) +tests/unit/pack-artifact-policy.test.tsdocs/ops/RELEASE_CHECKLIST.md— note the auto co-location🤖 Generated with Claude Code