docs: compression section — the 9-engine multi-layer stack - #3894
Conversation
Document the real pipeline (Session-Dedup, CCR, RTK, Headroom, Caveman, LLMLingua-2, Lite, Aggressive, Ultra) that replaced the old RTK+Caveman framing; presets table kept. '7 options' -> '9 engines'.
There was a problem hiding this comment.
Code Review
This pull request updates the README.md to document the new 9-engine compression stack, replacing the previous reference to 7 options. The reviewer noted that the listed order of the engines in the documentation table does not match the actual execution pipeline order defined in the codebase (src/lib/db/compressionCombos.ts) and provided a suggestion to reorder the table items to accurately reflect the execution priority.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| | # | Engine | What it does | | ||
| | --- | ----------------- | ------------------------------------------------------------------------ | | ||
| | 1 | **Session-Dedup** | Drops content repeated across turns (content-addressed, cross-turn) | | ||
| | 2 | **CCR** | Archives large blocks behind retrieve markers, fetched on demand | | ||
| | 3 | **RTK** | Smart tool-result filtering, dedup & truncation (command-aware) | | ||
| | 4 | **Headroom** | Lossless tabular compaction of homogeneous JSON arrays (~30%+) | | ||
| | 5 | **Caveman** | Rule-based prose compression (~65–75% on output) | | ||
| | 6 | **LLMLingua-2** | ML semantic pruning via MobileBERT ONNX — code-safe, async | | ||
| | 7 | **Lite** | Whitespace + image-URL trimming (latency-light baseline) | | ||
| | 8 | **Aggressive** | Summarization + progressive aging of old turns | | ||
| | 9 | **Ultra** | Heuristic token pruning with an optional small-model (SLM) tier | |
There was a problem hiding this comment.
The table lists the engines in an order that does not match the actual execution order (pipeline order) defined in the codebase. According to ENGINE_STACK_PRIORITY in src/lib/db/compressionCombos.ts, the engines are sorted by ascending priority:
session-dedup(3)ccr(4)lite(5)rtk(10)headroom(15)caveman(20)aggressive(30)llmlingua(35)ultra(40)
To accurately reflect the pipeline order, Lite should be listed as #3, Aggressive as #7, and LLMLingua-2 as #8.
| | # | Engine | What it does | | |
| | --- | ----------------- | ------------------------------------------------------------------------ | | |
| | 1 | **Session-Dedup** | Drops content repeated across turns (content-addressed, cross-turn) | | |
| | 2 | **CCR** | Archives large blocks behind retrieve markers, fetched on demand | | |
| | 3 | **RTK** | Smart tool-result filtering, dedup & truncation (command-aware) | | |
| | 4 | **Headroom** | Lossless tabular compaction of homogeneous JSON arrays (~30%+) | | |
| | 5 | **Caveman** | Rule-based prose compression (~65–75% on output) | | |
| | 6 | **LLMLingua-2** | ML semantic pruning via MobileBERT ONNX — code-safe, async | | |
| | 7 | **Lite** | Whitespace + image-URL trimming (latency-light baseline) | | |
| | 8 | **Aggressive** | Summarization + progressive aging of old turns | | |
| | 9 | **Ultra** | Heuristic token pruning with an optional small-model (SLM) tier | | |
| | # | Engine | What it does | | |
| | --- | ----------------- | ------------------------------------------------------------------------ | | |
| | 1 | **Session-Dedup** | Drops content repeated across turns (content-addressed, cross-turn) | | |
| | 2 | **CCR** | Archives large blocks behind retrieve markers, fetched on demand | | |
| | 3 | **Lite** | Whitespace + image-URL trimming (latency-light baseline) | | |
| | 4 | **RTK** | Smart tool-result filtering, dedup & truncation (command-aware) | | |
| | 5 | **Headroom** | Lossless tabular compaction of homogeneous JSON arrays (~30%+) | | |
| | 6 | **Caveman** | Rule-based prose compression (~65–75% on output) | | |
| | 7 | **Aggressive** | Summarization + progressive aging of old turns | | |
| | 8 | **LLMLingua-2** | ML semantic pruning via MobileBERT ONNX — code-safe, async | | |
| | 9 | **Ultra** | Heuristic token pruning with an optional small-model (SLM) tier | |
…iegosouzapw#3894) Document the real pipeline (Session-Dedup, CCR, RTK, Headroom, Caveman, LLMLingua-2, Lite, Aggressive, Ultra) that replaced the old RTK+Caveman framing; presets table kept. '7 options' -> '9 engines'.
Updates the README compression section to reflect the current system: a pipeline of 9 composable engines (was framed as just RTK + Caveman).
What changed
Docs-only change. The presentation site's compression section was updated in parallel (OmniRouteSite#1).