Skip to content

feat(compression): wire ultra's modelPath/slmFallbackToAggressive to the llmlingua SLM tier - #4257

Merged
diegosouzapw merged 1 commit into
release/v3.8.29from
feat/compression-ultra-slm-tier
Jun 19, 2026
Merged

diegosouzapw merged 1 commit into
release/v3.8.29from
feat/compression-ultra-slm-tier

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

Dá função real a modelPath / slmFallbackToAggressive do ultra (loose-end do #4253)

Contexto

Ao remover o seam SLM morto (#4253) sobrou a observação: o config do ultra ainda carregava modelPath / slmFallbackToAggressive, settáveis no schema mas inertes (o ultraCompress é heurística pura pruneByScore e ignorava ambos). Você pediu para ligar o engine llmlingua e dar função real.

O que faz

applyCompressionAsync agora roteia o modo ultra por uma camada SLM real quando modelPath está setado: a prosa é comprimida pelo engine llmlingua (compressor de modelo local, worker ONNX). O backend do llmlingua fail-opens quando o modelo não está provisionado, então degrada graciosamente:

Situação Resultado
modelo presente + ganho resultado do SLM (taggeado ultra-slm)
modelo ausente / sem ganho / falha + slmFallbackToAggressive aggressive
idem, sem slmFallbackToAggressive heurística ultra (pruneByScore)
sem modelPath heurística ultra (byte-idêntico ao atual)

Por que é seguro / não-invasivo

  • Default do ultra não define modelPath → uso existente do ultra inalterado.
  • Schema já aceitava os 2 campos (settáveis) → sem mudança de schema/DB.
  • Caminhos sync seguem heurística — inferência de modelo exige o entry async (onde a camada SLM vive, igual ao próprio llmlingua).
  • ⭐ Respeita o floor minTokens (2000) do llmlingua → prompts pequenos caem no fallback barato; prompts grandes usam o modelo.

Teste (TDD — RED 2→GREEN, backend injetável, sem modelo real)

tests/unit/compression/ultra-slm-tier.test.ts via setLlmlinguaBackend: SLM-roda, fallback-aggressive, fallback-heurística, e no-modelPath-não-toca-o-modelo.

Validação: suíte de compressão 713/713 ✓ · typecheck:core ✓ · lint ✓ · check:cycles ✓

Nota

O caminho de modelo real (ONNX ~99MB + optional-deps @atjsh/llmlingua-2) continua sendo provisioning de VPS — este PR torna o wiring funcional e fail-safe; com o modelo instalado, modelPath passa a comprimir de verdade.

…the llmlingua SLM tier

ultra was a pure heuristic (pruneByScore) and its modelPath / slmFallbackToAggressive
config fields were inert — settable in the schema but with no effect (the loose end
flagged when removing the dead SLM seam in #4253).

applyCompressionAsync now routes mode "ultra" through a real SLM tier when modelPath
is set: prose is compressed by the llmlingua engine (the local-model compressor). The
llmlingua backend fail-opens when the ONNX model is absent, so it degrades gracefully:
  - model present + gain → SLM result (tagged "ultra-slm");
  - model absent / no gain / failure → aggressive (when slmFallbackToAggressive),
    otherwise the heuristic ultra.

Without modelPath the behavior is byte-identical to the heuristic ultra (the default
ultra config sets no modelPath), so existing ultra usage is unchanged. Sync paths stay
heuristic — model inference requires the async entry point.

Tested via the injectable llmlingua backend (no real model needed): SLM-runs,
fallback-to-aggressive, fallback-to-heuristic, and no-modelPath-skips-the-model.
Full compression suite 713/713 green; typecheck:core + lint + check:cycles clean.
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@diegosouzapw
diegosouzapw merged commit a1a9f37 into release/v3.8.29 Jun 19, 2026
4 checks passed
@diegosouzapw
diegosouzapw deleted the feat/compression-ultra-slm-tier branch June 19, 2026 12:06
diegosouzapw added a commit that referenced this pull request Jun 19, 2026
…(postinstall) (#4286)

The compression "ultra" SLM tier (#4257) runs @atjsh/llmlingua-2 + transformers + tfjs
+ js-tiktoken in a worker thread shipped under dist/. These are optionalDependencies
installed into the ROOT node_modules on --include=optional, but the Next.js standalone
trace bundles ONLY @huggingface/transformers (3.5.2, pinned) into dist/node_modules —
not the dynamically-imported optionals.

Result: the worker resolves transformers from dist/node_modules (3.5.2) for its env
config but resolves @atjsh/llmlingua-2 from the ROOT, whose own transformers import
hits a DIFFERENT instance. The cacheDir config never reaches the instance llmlingua-2
uses, so the local model never loads and the SLM tier silently fails-open (and on a
root transformers 4.x, llmlingua-2 throws on the tokenizer API change).

Fix: postinstall co-locates the SLM optional closure from the root node_modules into
dist/node_modules (no-clobber, so the pinned dist transformers/onnxruntime stay), so
the worker resolves a single 3.5.2 instance and the local model loads.

VPS-validated (Rule #18): the co-located layout produced real 54.8% compression
(11520->5203 chars) via real ONNX inference on the production host, both the default
and the #4257 modelPath code paths.

- scripts/build/colocateOptionals.mjs: closure walk (deps+optionalDeps, skips the
  transformers peer) + no-clobber co-location; idempotent + fail-soft
- wired into scripts/build/postinstall.mjs next to ensureSwcHelpers
- registered in package.json files + pack-artifact allow/required lists
- tests/unit/colocate-optionals.test.ts: closure, no-clobber, idempotence, gates
- docs/ops/RELEASE_CHECKLIST.md: note the auto co-location
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
…the llmlingua SLM tier (diegosouzapw#4257)

Wire ultra's modelPath/slmFallbackToAggressive to the llmlingua SLM tier. Follow-up to diegosouzapw#4253. Integrated into release/v3.8.29.
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
…(postinstall) (diegosouzapw#4286)

The compression "ultra" SLM tier (diegosouzapw#4257) runs @atjsh/llmlingua-2 + transformers + tfjs
+ js-tiktoken in a worker thread shipped under dist/. These are optionalDependencies
installed into the ROOT node_modules on --include=optional, but the Next.js standalone
trace bundles ONLY @huggingface/transformers (3.5.2, pinned) into dist/node_modules —
not the dynamically-imported optionals.

Result: the worker resolves transformers from dist/node_modules (3.5.2) for its env
config but resolves @atjsh/llmlingua-2 from the ROOT, whose own transformers import
hits a DIFFERENT instance. The cacheDir config never reaches the instance llmlingua-2
uses, so the local model never loads and the SLM tier silently fails-open (and on a
root transformers 4.x, llmlingua-2 throws on the tokenizer API change).

Fix: postinstall co-locates the SLM optional closure from the root node_modules into
dist/node_modules (no-clobber, so the pinned dist transformers/onnxruntime stay), so
the worker resolves a single 3.5.2 instance and the local model loads.

VPS-validated (Rule diegosouzapw#18): the co-located layout produced real 54.8% compression
(11520->5203 chars) via real ONNX inference on the production host, both the default
and the diegosouzapw#4257 modelPath code paths.

- scripts/build/colocateOptionals.mjs: closure walk (deps+optionalDeps, skips the
  transformers peer) + no-clobber co-location; idempotent + fail-soft
- wired into scripts/build/postinstall.mjs next to ensureSwcHelpers
- registered in package.json files + pack-artifact allow/required lists
- tests/unit/colocate-optionals.test.ts: closure, no-clobber, idempotence, gates
- docs/ops/RELEASE_CHECKLIST.md: note the auto co-location
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant