feat(compression): wire ultra's modelPath/slmFallbackToAggressive to the llmlingua SLM tier - #4257
Merged
Conversation
…the llmlingua SLM tier ultra was a pure heuristic (pruneByScore) and its modelPath / slmFallbackToAggressive config fields were inert — settable in the schema but with no effect (the loose end flagged when removing the dead SLM seam in #4253). applyCompressionAsync now routes mode "ultra" through a real SLM tier when modelPath is set: prose is compressed by the llmlingua engine (the local-model compressor). The llmlingua backend fail-opens when the ONNX model is absent, so it degrades gracefully: - model present + gain → SLM result (tagged "ultra-slm"); - model absent / no gain / failure → aggressive (when slmFallbackToAggressive), otherwise the heuristic ultra. Without modelPath the behavior is byte-identical to the heuristic ultra (the default ultra config sets no modelPath), so existing ultra usage is unchanged. Sync paths stay heuristic — model inference requires the async entry point. Tested via the injectable llmlingua backend (no real model needed): SLM-runs, fallback-to-aggressive, fallback-to-heuristic, and no-modelPath-skips-the-model. Full compression suite 713/713 green; typecheck:core + lint + check:cycles clean.
Contributor
|
Warning You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again! |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
This was referenced Jun 19, 2026
Merged
diegosouzapw
added a commit
that referenced
this pull request
Jun 19, 2026
…(postinstall) (#4286) The compression "ultra" SLM tier (#4257) runs @atjsh/llmlingua-2 + transformers + tfjs + js-tiktoken in a worker thread shipped under dist/. These are optionalDependencies installed into the ROOT node_modules on --include=optional, but the Next.js standalone trace bundles ONLY @huggingface/transformers (3.5.2, pinned) into dist/node_modules — not the dynamically-imported optionals. Result: the worker resolves transformers from dist/node_modules (3.5.2) for its env config but resolves @atjsh/llmlingua-2 from the ROOT, whose own transformers import hits a DIFFERENT instance. The cacheDir config never reaches the instance llmlingua-2 uses, so the local model never loads and the SLM tier silently fails-open (and on a root transformers 4.x, llmlingua-2 throws on the tokenizer API change). Fix: postinstall co-locates the SLM optional closure from the root node_modules into dist/node_modules (no-clobber, so the pinned dist transformers/onnxruntime stay), so the worker resolves a single 3.5.2 instance and the local model loads. VPS-validated (Rule #18): the co-located layout produced real 54.8% compression (11520->5203 chars) via real ONNX inference on the production host, both the default and the #4257 modelPath code paths. - scripts/build/colocateOptionals.mjs: closure walk (deps+optionalDeps, skips the transformers peer) + no-clobber co-location; idempotent + fail-soft - wired into scripts/build/postinstall.mjs next to ensureSwcHelpers - registered in package.json files + pack-artifact allow/required lists - tests/unit/colocate-optionals.test.ts: closure, no-clobber, idempotence, gates - docs/ops/RELEASE_CHECKLIST.md: note the auto co-location
tkgo11
pushed a commit
to tkgo11/OmniRoute
that referenced
this pull request
Sep 23, 2026
…the llmlingua SLM tier (diegosouzapw#4257) Wire ultra's modelPath/slmFallbackToAggressive to the llmlingua SLM tier. Follow-up to diegosouzapw#4253. Integrated into release/v3.8.29.
tkgo11
pushed a commit
to tkgo11/OmniRoute
that referenced
this pull request
Sep 23, 2026
…(postinstall) (diegosouzapw#4286) The compression "ultra" SLM tier (diegosouzapw#4257) runs @atjsh/llmlingua-2 + transformers + tfjs + js-tiktoken in a worker thread shipped under dist/. These are optionalDependencies installed into the ROOT node_modules on --include=optional, but the Next.js standalone trace bundles ONLY @huggingface/transformers (3.5.2, pinned) into dist/node_modules — not the dynamically-imported optionals. Result: the worker resolves transformers from dist/node_modules (3.5.2) for its env config but resolves @atjsh/llmlingua-2 from the ROOT, whose own transformers import hits a DIFFERENT instance. The cacheDir config never reaches the instance llmlingua-2 uses, so the local model never loads and the SLM tier silently fails-open (and on a root transformers 4.x, llmlingua-2 throws on the tokenizer API change). Fix: postinstall co-locates the SLM optional closure from the root node_modules into dist/node_modules (no-clobber, so the pinned dist transformers/onnxruntime stay), so the worker resolves a single 3.5.2 instance and the local model loads. VPS-validated (Rule diegosouzapw#18): the co-located layout produced real 54.8% compression (11520->5203 chars) via real ONNX inference on the production host, both the default and the diegosouzapw#4257 modelPath code paths. - scripts/build/colocateOptionals.mjs: closure walk (deps+optionalDeps, skips the transformers peer) + no-clobber co-location; idempotent + fail-soft - wired into scripts/build/postinstall.mjs next to ensureSwcHelpers - registered in package.json files + pack-artifact allow/required lists - tests/unit/colocate-optionals.test.ts: closure, no-clobber, idempotence, gates - docs/ops/RELEASE_CHECKLIST.md: note the auto co-location
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Dá função real a
modelPath/slmFallbackToAggressivedo ultra (loose-end do #4253)Contexto
Ao remover o seam SLM morto (#4253) sobrou a observação: o config do ultra ainda carregava
modelPath/slmFallbackToAggressive, settáveis no schema mas inertes (oultraCompressé heurística purapruneByScoree ignorava ambos). Você pediu para ligar o engine llmlingua e dar função real.O que faz
applyCompressionAsyncagora roteia o modoultrapor uma camada SLM real quandomodelPathestá setado: a prosa é comprimida pelo engine llmlingua (compressor de modelo local, worker ONNX). O backend do llmlingua fail-opens quando o modelo não está provisionado, então degrada graciosamente:ultra-slm)slmFallbackToAggressiveslmFallbackToAggressivepruneByScore)modelPathPor que é seguro / não-invasivo
modelPath→ uso existente do ultra inalterado.minTokens(2000) do llmlingua → prompts pequenos caem no fallback barato; prompts grandes usam o modelo.Teste (TDD — RED 2→GREEN, backend injetável, sem modelo real)
tests/unit/compression/ultra-slm-tier.test.tsviasetLlmlinguaBackend: SLM-roda, fallback-aggressive, fallback-heurística, e no-modelPath-não-toca-o-modelo.Validação: suíte de compressão 713/713 ✓ ·
typecheck:core✓ ·lint✓ ·check:cycles✓Nota
O caminho de modelo real (ONNX ~99MB + optional-deps
@atjsh/llmlingua-2) continua sendo provisioning de VPS — este PR torna o wiring funcional e fail-safe; com o modelo instalado,modelPathpassa a comprimir de verdade.