Skip to content

docs: compression section — the 9-engine multi-layer stack - #3894

Merged
diegosouzapw merged 1 commit into
release/v3.8.26from
docs/compression-multilayer
Jun 15, 2026
Merged

diegosouzapw merged 1 commit into
release/v3.8.26from
docs/compression-multilayer

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

Updates the README compression section to reflect the current system: a pipeline of 9 composable engines (was framed as just RTK + Caveman).

What changed

  • New 9-engine stack table: Session-Dedup · CCR · RTK · Headroom (lossless tabular) · Caveman · LLMLingua-2 (ML/ONNX) · Lite · Aggressive · Ultra — in pipeline order, mix & match per combo.
  • Note that code/URLs/structured data are always preserved byte-perfect.
  • Presets table kept; diagram "7 options" → "9 engines".

Docs-only change. The presentation site's compression section was updated in parallel (OmniRouteSite#1).

Document the real pipeline (Session-Dedup, CCR, RTK, Headroom, Caveman,
LLMLingua-2, Lite, Aggressive, Ultra) that replaced the old RTK+Caveman
framing; presets table kept. '7 options' -> '9 engines'.
@diegosouzapw
diegosouzapw merged commit 7eff736 into release/v3.8.26 Jun 15, 2026
2 checks passed

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the README.md to document the new 9-engine compression stack, replacing the previous reference to 7 options. The reviewer noted that the listed order of the engines in the documentation table does not match the actual execution pipeline order defined in the codebase (src/lib/db/compressionCombos.ts) and provided a suggestion to reorder the table items to accurately reflect the execution priority.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread README.md
Comment on lines +448 to +458
| # | Engine | What it does |
| --- | ----------------- | ------------------------------------------------------------------------ |
| 1 | **Session-Dedup** | Drops content repeated across turns (content-addressed, cross-turn) |
| 2 | **CCR** | Archives large blocks behind retrieve markers, fetched on demand |
| 3 | **RTK** | Smart tool-result filtering, dedup & truncation (command-aware) |
| 4 | **Headroom** | Lossless tabular compaction of homogeneous JSON arrays (~30%+) |
| 5 | **Caveman** | Rule-based prose compression (~65–75% on output) |
| 6 | **LLMLingua-2** | ML semantic pruning via MobileBERT ONNX — code-safe, async |
| 7 | **Lite** | Whitespace + image-URL trimming (latency-light baseline) |
| 8 | **Aggressive** | Summarization + progressive aging of old turns |
| 9 | **Ultra** | Heuristic token pruning with an optional small-model (SLM) tier |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The table lists the engines in an order that does not match the actual execution order (pipeline order) defined in the codebase. According to ENGINE_STACK_PRIORITY in src/lib/db/compressionCombos.ts, the engines are sorted by ascending priority:

  • session-dedup (3)
  • ccr (4)
  • lite (5)
  • rtk (10)
  • headroom (15)
  • caveman (20)
  • aggressive (30)
  • llmlingua (35)
  • ultra (40)

To accurately reflect the pipeline order, Lite should be listed as #3, Aggressive as #7, and LLMLingua-2 as #8.

Suggested change
| # | Engine | What it does |
| --- | ----------------- | ------------------------------------------------------------------------ |
| 1 | **Session-Dedup** | Drops content repeated across turns (content-addressed, cross-turn) |
| 2 | **CCR** | Archives large blocks behind retrieve markers, fetched on demand |
| 3 | **RTK** | Smart tool-result filtering, dedup & truncation (command-aware) |
| 4 | **Headroom** | Lossless tabular compaction of homogeneous JSON arrays (~30%+) |
| 5 | **Caveman** | Rule-based prose compression (~65–75% on output) |
| 6 | **LLMLingua-2** | ML semantic pruning via MobileBERT ONNX — code-safe, async |
| 7 | **Lite** | Whitespace + image-URL trimming (latency-light baseline) |
| 8 | **Aggressive** | Summarization + progressive aging of old turns |
| 9 | **Ultra** | Heuristic token pruning with an optional small-model (SLM) tier |
| # | Engine | What it does |
| --- | ----------------- | ------------------------------------------------------------------------ |
| 1 | **Session-Dedup** | Drops content repeated across turns (content-addressed, cross-turn) |
| 2 | **CCR** | Archives large blocks behind retrieve markers, fetched on demand |
| 3 | **Lite** | Whitespace + image-URL trimming (latency-light baseline) |
| 4 | **RTK** | Smart tool-result filtering, dedup & truncation (command-aware) |
| 5 | **Headroom** | Lossless tabular compaction of homogeneous JSON arrays (~30%+) |
| 6 | **Caveman** | Rule-based prose compression (~65–75% on output) |
| 7 | **Aggressive** | Summarization + progressive aging of old turns |
| 8 | **LLMLingua-2** | ML semantic pruning via MobileBERT ONNX — code-safe, async |
| 9 | **Ultra** | Heuristic token pruning with an optional small-model (SLM) tier |

@diegosouzapw
diegosouzapw deleted the docs/compression-multilayer branch June 15, 2026 12:43
@diegosouzapw diegosouzapw mentioned this pull request Jun 16, 2026
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
…iegosouzapw#3894)

Document the real pipeline (Session-Dedup, CCR, RTK, Headroom, Caveman,
LLMLingua-2, Lite, Aggressive, Ultra) that replaced the old RTK+Caveman
framing; presets table kept. '7 options' -> '9 engines'.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant